Skip to content

Classification Logic

Stickler classifies every field comparison into one of five confusion-matrix categories. These categories drive all derived metrics (precision, recall, F1) and aggregate reporting.

Core Definitions

Category Abbr. Definition
True Positive TP GT and EST are both non-null and match above the threshold
False Alarm FA GT is null, EST is non-null
True Negative TN GT and EST are both null
False Negative FN GT is non-null, EST is null
False Discovery FD GT and EST are both non-null but match below the threshold

False Alarm (FA) and False Discovery (FD) together make up the broader False Positive (FP) count:

FP = FA + FD

Classification by Data Type

Simple Values (Strings, Numbers, Booleans)

Ground Truth Prediction Classification Notes
"value" "value" TP Exact match
"value" "similar" FD Both non-null, below threshold
"value" null FN Missing prediction
null "value" FA Spurious prediction
null null TN Correctly absent
"" null TN Empty string treated as null

Lists

Lists use the Hungarian algorithm for optimal element pairing.

  1. Empty lists: [] vs [] is TN; [] vs [items] produces one FA per item; [items] vs [] produces one FN per item.
  2. Matched elements: similarity >= threshold is TP; below threshold is FD.
  3. Unmatched elements: leftover GT elements are FN; leftover EST elements are FA.

Example: Mixed Matching

GT  = ["red", "blue", "green"]
EST = ["red", "yellow", "orange", "blue"]

The assignment pairs as many elements as the shorter list allows, so it makes three pairs here and leaves one prediction over:

GT EST Score Classification
"red" "red" 1.00 TP
"blue" "blue" 1.00 TP
"green" "orange" 0.17 FD
"yellow" FA

Result: TP=2, FD=1, FA=1, FN=0, and so FP=2

Note that "green" is an FD and not an FN. It did get a partner, and a low score is what the FD category is for. Only an element the assignment could not pair at all is FN or FA, which is why FN is zero whenever GT is the shorter list.

Example: Below-Threshold Matches

GT  = ["apple", "banana", "cherry"]
EST = ["appx", "bnn", "chry"]        (threshold = 0.7)

All pairs match below 0.7 -- each is FD.

Result: TP=0, FA=0, FN=0, FD=3

Nested Objects

Nested objects are evaluated recursively, field by field.

Condition Classification
Both have the field, similarity >= threshold TP
Both have the field, similarity < threshold FD
Only GT has the field FN
Only EST has the field FA

Example

GT  = {name: "John", age: 30, address: "123 Main St"}
EST = {name: "John", age: 31, phone: "555-1234"}
  • name: exact match -- TP
  • age: both present, mismatch -- FD
  • address: only in GT -- FN
  • phone: only in EST -- FA

Result: TP=1, FA=1, FN=1, FD=1

Objects of different classes

Two objects of different classes are always an FD, whatever their attributes say. Identical field names and identical values do not make them a match:

Pet(name="rex")  vs  Cat(name="rex")   ->  FD

This is the one FD in the table above that is not a threshold decision. Every other row compares a similarity against a cutoff; this one is settled before any comparator runs, because the class is part of the value's identity rather than metadata about it. A Cat is not a Pet that happens to score well.

The rule uses the EXACT class, so a subclass against its base is also an FD. A subclass carries fields the base does not, so the two do not describe the same shape, and scoring them on their shared fields would report a near-match for what is really a schema difference.

Stickler warns once per field rather than raising. Which class arrives is a property of the prediction, so raising would end a bulk run on document N after succeeding on N-1.

A correctly annotated field never reaches this rule: pydantic refuses a Dog for an Optional[Cat] field when the object is constructed. It applies where the annotation permitted more than one class (Union[Cat, Dog], Any, object) or where a subclass arrived for its base, which Optional[Pet] accepts.

Enforced for plain BaseModel objects, not yet for StructuredModel

As of this release the rule is applied to plain pydantic.BaseModel objects and to the elements of a list of them. Two StructuredModel objects of different classes are still scored field by field and can report TP:

Pet(name="rex") vs Cat(name="rex")     plain BaseModel    0.0   FD
                                       StructuredModel    1.0   TP

That is the older behaviour rather than a deliberate exception, and closing it is tracked in #327. Until then, treat a heterogeneous StructuredModel pair as unverified rather than as endorsed: annotate the field with a single model type if you need the guarantee today.

Whether a refused object should still pair inside a list, or should instead be counted as a missed ground truth plus an invented prediction, is the open question in #321.

Derived Metrics

From the base counts:

Metric Formula Meaning
Precision TP / (TP + FP) Fraction of predictions that are correct
Recall TP / (TP + FN) Fraction of ground-truth values found (FD in neither term; see below)
F1 Score 2 * Precision * Recall / (Precision + Recall) Harmonic mean of precision and recall
Accuracy (TP + TN) / (TP + TN + FP + FN) Overall correctness

An FD is invisible to recall under that formula, because it appears in neither term. The mixed-matching example above scores recall 1.000 even though "green" was never found, since it is an FD and not an FN. Set recall_with_fd=True for TP / (TP + FN + FD), which scores the same example 0.667. HungarianMatcher.calculate_metrics always uses that second formula. See FD and recall.

These are computed at every node of the result tree, from that node's own counts, and which counts those are depends on which node you read. overall classifies that node's direct children: for a list field, whether each pairing was genuine or spurious; at the root, its own fields, which means the two units can mix in a single count. aggregate gives leaf detail, and for a list of objects that means only the items that were comparable, since a list item below the element class's match_threshold is one FD and is not descended into. A single nested StructuredModel field is not gated this way: its leaves are always reported. See Which node answers which question.

Edge Cases

Null vs. empty equivalence -- Empty strings (""), empty lists ([]), and empty objects ({}) are treated as null. Comparing any of these with null yields TN.

TN in object matching -- A TN is absence of evidence, not evidence of a match. Fields absent on both sides are excluded from the weighted object similarity used by Hungarian matching. If every field is absent, the similarity is defined as 1.0.

Threshold boundary -- A similarity score exactly equal to the threshold counts as a match (TP).

List order -- Order does not matter. The Hungarian algorithm finds the optimal pairing regardless of element position.

Nested lists -- For List[StructuredModel], the Hungarian algorithm pairs objects at the list level, then each matched pair is evaluated recursively.

Missing vs. null fields -- A missing field and a field explicitly set to null are handled the same way: if the other side has a non-null value, the result is FN or FA accordingly.

See Also