Add high-level dnn text Model API (TextRecognition + TextDetection EAST/DB) - #1918
Merged
Merged
Conversation
…ST/DB) Follow-up to the dnn Model API work (#1917). Adds the remaining high-level text models from OpenCV 5's dnn module. - TextRecognitionModel: SetDecodeType/GetDecodeType, SetDecodeOptsCTCPrefixBeamSearch, SetVocabulary/GetVocabulary, Recognize. - TextDetectionModel (abstract base): Detect (quadrangles + confidences) and DetectTextRectangles (RotatedRects + confidences), shared by EAST and DB. - TextDetectionModelEAST: SetConfidenceThreshold/Get, SetNMSThreshold/Get. - TextDetectionModelDB: SetBinaryThreshold/Get, SetPolygonThreshold/Get, SetUnclipRatio/Get, SetMaxCandidates/Get. All classes can be built from a file path or from a Net, and inherit the base Model preprocessing setters. As with the Model API, the base TextDetectionModel methods upcast the derived pointer (single inheritance, no pointer adjustment). Native dnn_TextModel.h is included from dnn.cpp and registered in the vcxproj and .filters. Verified locally against a freshly built OpenCvSharpExtern (OpenCV 5.0.0): - The 3 lifecycle tests pass (constructors, setters/getters, vocabulary and decode-type round-trips for all three model types). - The EAST end-to-end test (temporarily un-skipped) downloads the EAST model and detects text rectangles on abbey_road.jpg with confidences above the threshold; it ships as [ExplicitTheory] like the existing EAST test. - Full dnn suite stays green (25 passed / 4 explicit-skipped). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This was referenced Jun 23, 2026
Unify free-function facades: nested Cv2.Dnn/Cv2.Aruco/... mirroring cv::dnn:: (Python cv2.dnn)
#1921
Closed
hundong2
pushed a commit
to hundong2/opencvsharp
that referenced
this pull request
Jul 4, 2026
…ed/soft NMS Priority group B of the dnn coverage work (follows the Model API in shimat#1917/shimat#1918). Adds the OpenCV 5 preprocessing / post-processing helpers that were missing. - readNetFromTFLite: CvDnn.ReadNetFromTFLite and Net.ReadNetFromTFLite (file path, byte[], ReadOnlySpan<byte>, Stream) mirroring the ONNX readers. - blobFromImageWithParams / blobFromImagesWithParams via a new Image2BlobParams class (ScaleFactor, Size, Mean, SwapRB, Depth, DataLayout, PaddingMode, BorderValue), enabling letterbox/center-crop preprocessing. - imagesFromBlob: CvDnn.ImagesFromBlob to unpack a 4D blob back into images. - NMSBoxesBatched (Rect and Rect2d) and softNMSBoxes (Rect). - New enums DataLayout, ImagePaddingMode, SoftNMSMethod. Note: in OpenCV 5 the DataLayout enum moved from cv::dnn to the cv namespace (core/mat.hpp); the native wrapper uses cv::DataLayout accordingly. Verified locally against a freshly built OpenCvSharpExtern (OpenCV 5.0.0); the 6 BlobAndNmsTest cases pass and the full dnn suite stays green: - blobFromImageWithParams produces the same blob as blobFromImage for matching params, and LETTERBOX yields the requested NCHW shape. - imagesFromBlob round-trips blobFromImages. - NMSBoxesBatched returns the expected kept indices; softNMSBoxes returns indices with matching updated scores. - readNetFromTFLite on an invalid buffer raises OpenCVException (entry point wired), rather than a P/Invoke failure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
hundong2
pushed a commit
to hundong2/opencvsharp
that referenced
this pull request
Jul 4, 2026
…e*, model format Priority group C/D of the dnn coverage work (follows shimat#1917/shimat#1918/shimat#1919). Wraps the genuinely useful and C#-friendly Net introspection / utility members and skips the ones that only make sense in C++. Net: - SetInputShape(name, shape) - GetParam / SetParam (by layer id and by layer name) - GetLayerTypes, GetLayersCount - GetModelFormat (+ new ModelFormat enum) - EnableWinograd, DumpToPbtxt, PrintPerfProfile - EnableKVCache / DisableKVCache / ResetKVCache (transformer/LLM KV cache control) CvDnn (free functions): - GetAvailableTargets(Backend), GetAvailableBackends(), EnableModelDiagnostics(bool) Deliberately NOT wrapped (documented in docs/opencv5/dnn-coverage.md): - getFLOPS / getLayersShapes / getLayerShapes / getMemoryConsumption — awkward in 5.0 (MatShape is now a struct, plus netInputTypes and nested vector<vector<MatShape>>); low value, can be added later. - Layer class, getLayer/addLayer/registerOutput/getLayerInputs — these exist mainly for authoring custom layers, which requires subclassing cv::dnn::Layer in C++ and is not practical via P/Invoke (per-forward managed callbacks + blob marshaling). - forwardAsync/AsyncArray — async-backend only, low value. Notes: - MatShape moved from std::vector<int> alias to a real cv::MatShape struct in OpenCV 5; SetInputShape constructs it from an int range natively. Verified locally against a freshly built OpenCvSharpExtern (OpenCV 5.0.0): 9 NetIntrospectionTest cases pass, full dnn suite green (31 passed / 3 explicit-skipped). GetModelFormat returns TF for a TF model; GetLayerTypes/GetLayersCount, GetAvailableTargets/Backends return real values. GetParam/SetParam, SetInputShape and DumpToPbtxt are verified via their error path (the committed TF models do not expose layer blobs, and DumpToPbtxt needs allocated output blobs), confirming the entry points are wired rather than failing in the P/Invoke layer. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Follow-up to #1917. Completes priority group A (high-level Model API) by adding
the remaining text models from OpenCV 5's dnn module.
What's added
src/OpenCvSharpExtern/dnn_TextModel.h(included fromdnn.cpp, registered in vcxproj/filters)NativeMethods_dnn_TextModel.csTextRecognitionModel.cs,TextDetectionModel.cs(abstract base),TextDetectionModelEAST.cs,TextDetectionModelDB.csdnn/TextModelTest.csAPI
TextRecognitionModel:SetDecodeType/GetDecodeType,SetDecodeOptsCTCPrefixBeamSearch,SetVocabulary/GetVocabulary,RecognizeTextDetectionModel(abstract base):Detect(quadrangles + confidences) andDetectTextRectangles(RotatedRects + confidences) — shared by EAST and DBTextDetectionModelEAST:SetConfidenceThreshold/Get,SetNMSThreshold/GetTextDetectionModelDB:SetBinaryThreshold/Get,SetPolygonThreshold/Get,SetUnclipRatio/Get,SetMaxCandidates/GetAll classes build from a file path or from a
Netand inherit the baseModelpreprocessing setters. The base
TextDetectionModelmethods upcast the derivedpointer to
cv::dnn::TextDetectionModel*(single inheritance, no pointer adjustment).Verification
Built
OpenCvSharpExternlocally against OpenCV 5.0.0 and ran the suite:decode-type round-trips for all three model types.
model and detects text rectangles on
abbey_road.jpgwith confidences above thethreshold. It ships as
[ExplicitTheory], matching the existingEastTextDetectionTest.Notes
TextRecognitionModel.recognizebatch overload (withroiRects) is intentionallyomitted for now —
InputArrayOfArraysmarshaling is awkward and the overload israrely used. Can be added later if needed.
Next
With group A complete, the next increment is priority B (
readNetFromTFLite,blobFromImageWithParams,NMSBoxesBatched/softNMSBoxes,imagesFromBlob).🤖 Generated with Claude Code