Skip to content

Add high-level dnn text Model API (TextRecognition + TextDetection EAST/DB) - #1918

Merged
shimat merged 1 commit into
5.xfrom
feature/dnn-text-models
Jun 23, 2026
Merged

shimat merged 1 commit into
5.xfrom
feature/dnn-text-models

Conversation

@shimat

@shimat shimat commented Jun 23, 2026

Copy link
Copy Markdown
Owner

Summary

Follow-up to #1917. Completes priority group A (high-level Model API) by adding
the remaining text models from OpenCV 5's dnn module.

What's added

Layer File
Native C++ src/OpenCvSharpExtern/dnn_TextModel.h (included from dnn.cpp, registered in vcxproj/filters)
P/Invoke NativeMethods_dnn_TextModel.cs
Managed wrappers TextRecognitionModel.cs, TextDetectionModel.cs (abstract base), TextDetectionModelEAST.cs, TextDetectionModelDB.cs
Tests dnn/TextModelTest.cs

API

  • TextRecognitionModel: SetDecodeType/GetDecodeType,
    SetDecodeOptsCTCPrefixBeamSearch, SetVocabulary/GetVocabulary, Recognize
  • TextDetectionModel (abstract base): Detect (quadrangles + confidences) and
    DetectTextRectangles (RotatedRects + confidences) — shared by EAST and DB
  • TextDetectionModelEAST: SetConfidenceThreshold/Get, SetNMSThreshold/Get
  • TextDetectionModelDB: SetBinaryThreshold/Get, SetPolygonThreshold/Get,
    SetUnclipRatio/Get, SetMaxCandidates/Get

All classes build from a file path or from a Net and inherit the base Model
preprocessing setters. The base TextDetectionModel methods upcast the derived
pointer to cv::dnn::TextDetectionModel* (single inheritance, no pointer adjustment).

Verification

Built OpenCvSharpExtern locally against OpenCV 5.0.0 and ran the suite:

  • 3 lifecycle tests pass — constructors, setters/getters, and vocabulary /
    decode-type round-trips for all three model types.
  • The EAST end-to-end test (temporarily un-skipped for the run) downloads the EAST
    model and detects text rectangles on abbey_road.jpg with confidences above the
    threshold. It ships as [ExplicitTheory], matching the existing EastTextDetectionTest.
  • Full dnn suite stays green (25 passed / 4 explicit-skipped).

Notes

  • TextRecognitionModel.recognize batch overload (with roiRects) is intentionally
    omitted for now — InputArrayOfArrays marshaling is awkward and the overload is
    rarely used. Can be added later if needed.

Next

With group A complete, the next increment is priority B (readNetFromTFLite,
blobFromImageWithParams, NMSBoxesBatched/softNMSBoxes, imagesFromBlob).

🤖 Generated with Claude Code

…ST/DB)

Follow-up to the dnn Model API work (#1917). Adds the remaining high-level
text models from OpenCV 5's dnn module.

- TextRecognitionModel: SetDecodeType/GetDecodeType,
  SetDecodeOptsCTCPrefixBeamSearch, SetVocabulary/GetVocabulary, Recognize.
- TextDetectionModel (abstract base): Detect (quadrangles + confidences) and
  DetectTextRectangles (RotatedRects + confidences), shared by EAST and DB.
- TextDetectionModelEAST: SetConfidenceThreshold/Get, SetNMSThreshold/Get.
- TextDetectionModelDB: SetBinaryThreshold/Get, SetPolygonThreshold/Get,
  SetUnclipRatio/Get, SetMaxCandidates/Get.

All classes can be built from a file path or from a Net, and inherit the base
Model preprocessing setters. As with the Model API, the base TextDetectionModel
methods upcast the derived pointer (single inheritance, no pointer adjustment).

Native dnn_TextModel.h is included from dnn.cpp and registered in the vcxproj
and .filters.

Verified locally against a freshly built OpenCvSharpExtern (OpenCV 5.0.0):
- The 3 lifecycle tests pass (constructors, setters/getters, vocabulary and
  decode-type round-trips for all three model types).
- The EAST end-to-end test (temporarily un-skipped) downloads the EAST model
  and detects text rectangles on abbey_road.jpg with confidences above the
  threshold; it ships as [ExplicitTheory] like the existing EAST test.
- Full dnn suite stays green (25 passed / 4 explicit-skipped).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@shimat shimat self-assigned this Jun 23, 2026
@shimat
shimat merged commit 42de9a5 into 5.x Jun 23, 2026
11 checks passed
@shimat
shimat deleted the feature/dnn-text-models branch June 23, 2026 13:12
hundong2 pushed a commit to hundong2/opencvsharp that referenced this pull request Jul 4, 2026
…ed/soft NMS

Priority group B of the dnn coverage work (follows the Model API in shimat#1917/shimat#1918).
Adds the OpenCV 5 preprocessing / post-processing helpers that were missing.

- readNetFromTFLite: CvDnn.ReadNetFromTFLite and Net.ReadNetFromTFLite
  (file path, byte[], ReadOnlySpan<byte>, Stream) mirroring the ONNX readers.
- blobFromImageWithParams / blobFromImagesWithParams via a new Image2BlobParams
  class (ScaleFactor, Size, Mean, SwapRB, Depth, DataLayout, PaddingMode,
  BorderValue), enabling letterbox/center-crop preprocessing.
- imagesFromBlob: CvDnn.ImagesFromBlob to unpack a 4D blob back into images.
- NMSBoxesBatched (Rect and Rect2d) and softNMSBoxes (Rect).
- New enums DataLayout, ImagePaddingMode, SoftNMSMethod.

Note: in OpenCV 5 the DataLayout enum moved from cv::dnn to the cv namespace
(core/mat.hpp); the native wrapper uses cv::DataLayout accordingly.

Verified locally against a freshly built OpenCvSharpExtern (OpenCV 5.0.0); the
6 BlobAndNmsTest cases pass and the full dnn suite stays green:
- blobFromImageWithParams produces the same blob as blobFromImage for matching
  params, and LETTERBOX yields the requested NCHW shape.
- imagesFromBlob round-trips blobFromImages.
- NMSBoxesBatched returns the expected kept indices; softNMSBoxes returns
  indices with matching updated scores.
- readNetFromTFLite on an invalid buffer raises OpenCVException (entry point
  wired), rather than a P/Invoke failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
hundong2 pushed a commit to hundong2/opencvsharp that referenced this pull request Jul 4, 2026
…e*, model format

Priority group C/D of the dnn coverage work (follows shimat#1917/shimat#1918/shimat#1919). Wraps the
genuinely useful and C#-friendly Net introspection / utility members and skips the
ones that only make sense in C++.

Net:
- SetInputShape(name, shape)
- GetParam / SetParam (by layer id and by layer name)
- GetLayerTypes, GetLayersCount
- GetModelFormat (+ new ModelFormat enum)
- EnableWinograd, DumpToPbtxt, PrintPerfProfile
- EnableKVCache / DisableKVCache / ResetKVCache (transformer/LLM KV cache control)

CvDnn (free functions):
- GetAvailableTargets(Backend), GetAvailableBackends(), EnableModelDiagnostics(bool)

Deliberately NOT wrapped (documented in docs/opencv5/dnn-coverage.md):
- getFLOPS / getLayersShapes / getLayerShapes / getMemoryConsumption — awkward in 5.0
  (MatShape is now a struct, plus netInputTypes and nested vector<vector<MatShape>>);
  low value, can be added later.
- Layer class, getLayer/addLayer/registerOutput/getLayerInputs — these exist mainly for
  authoring custom layers, which requires subclassing cv::dnn::Layer in C++ and is not
  practical via P/Invoke (per-forward managed callbacks + blob marshaling).
- forwardAsync/AsyncArray — async-backend only, low value.

Notes:
- MatShape moved from std::vector<int> alias to a real cv::MatShape struct in OpenCV 5;
  SetInputShape constructs it from an int range natively.

Verified locally against a freshly built OpenCvSharpExtern (OpenCV 5.0.0): 9
NetIntrospectionTest cases pass, full dnn suite green (31 passed / 3 explicit-skipped).
GetModelFormat returns TF for a TF model; GetLayerTypes/GetLayersCount,
GetAvailableTargets/Backends return real values. GetParam/SetParam, SetInputShape and
DumpToPbtxt are verified via their error path (the committed TF models do not expose
layer blobs, and DumpToPbtxt needs allocated output blobs), confirming the entry points
are wired rather than failing in the P/Invoke layer.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@shimat shimat added the enhancement New feature or improvement to OpenCvSharp label Jul 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or improvement to OpenCvSharp

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant