Classification Logic
Stickler classifies every field comparison into one of five confusion-matrix categories. These categories drive all derived metrics (precision, recall, F1) and aggregate reporting.
Core Definitions
| Category | Abbr. | Definition |
|---|---|---|
| True Positive | TP | GT and EST are both non-null and match above the threshold |
| False Alarm | FA | GT is null, EST is non-null |
| True Negative | TN | GT and EST are both null |
| False Negative | FN | GT is non-null, EST is null |
| False Discovery | FD | GT and EST are both non-null but match below the threshold |
False Alarm (FA) and False Discovery (FD) together make up the broader False Positive (FP) count:
FP = FA + FD
Classification by Data Type
Simple Values (Strings, Numbers, Booleans)
| Ground Truth | Prediction | Classification | Notes |
|---|---|---|---|
"value" |
"value" |
TP | Exact match |
"value" |
"similar" |
FD | Both non-null, below threshold |
"value" |
null |
FN | Missing prediction |
null |
"value" |
FA | Spurious prediction |
null |
null |
TN | Correctly absent |
"" |
null |
TN | Empty string treated as null |
Lists
Lists use the Hungarian algorithm for optimal element pairing.
- Empty lists:
[] vs []is TN;[] vs [items]produces one FA per item;[items] vs []produces one FN per item. - Matched elements: similarity >= threshold is TP; below threshold is FD.
- Unmatched elements: leftover GT elements are FN; leftover EST elements are FA.
Example: Mixed Matching
GT = ["red", "blue", "green"]
EST = ["red", "yellow", "orange", "blue"]
The assignment pairs as many elements as the shorter list allows, so it makes three pairs here and leaves one prediction over:
| GT | EST | Score | Classification |
|---|---|---|---|
"red" |
"red" |
1.00 | TP |
"blue" |
"blue" |
1.00 | TP |
"green" |
"orange" |
0.17 | FD |
"yellow" |
FA |
Result: TP=2, FD=1, FA=1, FN=0, and so FP=2
Note that "green" is an FD and not an FN. It did get a partner, and a low
score is what the FD category is for. Only an element the assignment could not
pair at all is FN or FA, which is why FN is zero whenever GT is the shorter
list.
Example: Below-Threshold Matches
GT = ["apple", "banana", "cherry"]
EST = ["appx", "bnn", "chry"] (threshold = 0.7)
All pairs match below 0.7 -- each is FD.
Result: TP=0, FA=0, FN=0, FD=3
Nested Objects
Nested objects are evaluated recursively, field by field.
| Condition | Classification |
|---|---|
| Both have the field, similarity >= threshold | TP |
| Both have the field, similarity < threshold | FD |
| Only GT has the field | FN |
| Only EST has the field | FA |
Example
GT = {name: "John", age: 30, address: "123 Main St"}
EST = {name: "John", age: 31, phone: "555-1234"}
name: exact match -- TPage: both present, mismatch -- FDaddress: only in GT -- FNphone: only in EST -- FA
Result: TP=1, FA=1, FN=1, FD=1
Objects of different classes
Two objects of different classes are always an FD, whatever their attributes say. Identical field names and identical values do not make them a match:
Pet(name="rex") vs Cat(name="rex") -> FD
This is the one FD in the table above that is not a threshold decision. Every
other row compares a similarity against a cutoff; this one is settled before any
comparator runs, because the class is part of the value's identity rather than
metadata about it. A Cat is not a Pet that happens to score well.
The rule uses the EXACT class, so a subclass against its base is also an FD. A subclass carries fields the base does not, so the two do not describe the same shape, and scoring them on their shared fields would report a near-match for what is really a schema difference.
Stickler warns once per field rather than raising. Which class arrives is a property of the prediction, so raising would end a bulk run on document N after succeeding on N-1.
A correctly annotated field never reaches this rule: pydantic refuses a Dog for
an Optional[Cat] field when the object is constructed. It applies where the
annotation permitted more than one class (Union[Cat, Dog], Any, object) or
where a subclass arrived for its base, which Optional[Pet] accepts.
Enforced for plain BaseModel objects, not yet for StructuredModel
As of this release the rule is applied to plain pydantic.BaseModel objects
and to the elements of a list of them. Two StructuredModel objects of
different classes are still scored field by field and can report TP:
Pet(name="rex") vs Cat(name="rex") plain BaseModel 0.0 FD
StructuredModel 1.0 TP
That is the older behaviour rather than a deliberate exception, and closing it is
tracked in #327. Until then, treat a heterogeneous StructuredModel pair as
unverified rather than as endorsed: annotate the field with a single model type
if you need the guarantee today.
Whether a refused object should still pair inside a list, or should instead be counted as a missed ground truth plus an invented prediction, is the open question in #321.
Derived Metrics
From the base counts:
| Metric | Formula | Meaning |
|---|---|---|
| Precision | TP / (TP + FP) | Fraction of predictions that are correct |
| Recall | TP / (TP + FN) | Fraction of ground-truth values found (FD in neither term; see below) |
| F1 Score | 2 * Precision * Recall / (Precision + Recall) | Harmonic mean of precision and recall |
| Accuracy | (TP + TN) / (TP + TN + FP + FN) | Overall correctness |
An FD is invisible to recall under that formula, because it appears in
neither term. The mixed-matching example above scores recall 1.000 even though
"green" was never found, since it is an FD and not an FN. Set
recall_with_fd=True for TP / (TP + FN + FD), which scores the same example
0.667. HungarianMatcher.calculate_metrics always uses that second formula.
See FD and recall.
These are computed at every node of the result tree, from that node's own counts, and which counts those are depends on which node you read. overall classifies that node's direct children: for a list field, whether each pairing was genuine or spurious; at the root, its own fields, which means the two units can mix in a single count. aggregate gives leaf detail, and for a list of objects that means only the items that were comparable, since a list item below the element class's match_threshold is one FD and is not descended into. A single nested StructuredModel field is not gated this way: its leaves are always reported. See Which node answers which question.
Edge Cases
Null vs. empty equivalence -- Empty strings (""), empty lists ([]), and empty objects ({}) are treated as null. Comparing any of these with null yields TN.
TN in object matching -- A TN is absence of evidence, not evidence of a match. Fields absent on both sides are excluded from the weighted object similarity used by Hungarian matching. If every field is absent, the similarity is defined as 1.0.
Threshold boundary -- A similarity score exactly equal to the threshold counts as a match (TP).
List order -- Order does not matter. The Hungarian algorithm finds the optimal pairing regardless of element position.
Nested lists -- For List[StructuredModel], the Hungarian algorithm pairs objects at the list level, then each matched pair is evaluated recursively.
Missing vs. null fields -- A missing field and a field explicitly set to null are handled the same way: if the other side has a non-null value, the result is FN or FA accordingly.
See Also
- Understanding Results -- interpreting the full result dictionary
- Hungarian Matching -- details on the list-pairing algorithm