Benchmarks
Benchmarks read saved JSONs and run locally. Constructing or evaluating a
benchmark does not call an LLM. See the stage-specific benchmarking guides for
input formats.
Classification and rotation
ClassificationBenchmark
dataclass
ClassificationBenchmark(*, true_path: Path, pred_path: Path)
Bases: Generic[LabelType]
Compare saved image labels with ground truth and report accuracy.
The benchmark reads JSON files without opening images or making model
requests. Binary, multiclass and rotation benchmarks inherit the shared
result properties documented here.
| Parameters: |
-
true_path
(Path)
–
Ground-truth JSON file containing a list of objects with exactly
img_path and label keys. Image paths must be strings; labels must
be str, int or bool values.
-
pred_path
(Path)
–
Prediction JSON file with the same format and number of entries as
true_path. Entries are paired by list position. Paired image
filenames, including extensions, must match; parent directories may
differ. The benchmark does not sort or reorder either list.
|
| Attributes: |
-
result_list
(tuple[SingleClassificationResult[LabelType], ...])
–
One record per paired image, in the JSON list order. Each record
contains img_path from ground truth, true and pred.
-
correct_predictions
(tuple[SingleClassificationResult[LabelType], ...])
–
Records whose predicted labels equal their ground-truth labels.
-
incorrect_predictions
(tuple[SingleClassificationResult[LabelType], ...])
–
Records whose predicted labels differ from their ground-truth labels.
-
accuracy
(float)
–
Number of correct predictions divided by the number of paired images.
Returns 0.0 for two empty JSON lists.
-
img_paths
(tuple[str, ...])
–
Image paths from the ground-truth JSON, in the JSON list order.
-
trues
(tuple[LabelType, ...])
–
Ground-truth labels, in the JSON list order.
-
preds
(tuple[LabelType, ...])
–
Predicted labels, in the JSON list order.
|
| Raises: |
-
OSError
–
If either JSON file cannot be opened.
-
TypeError
–
If a JSON document is not a list, an entry is not an object, an
img_path is not a string, or a label has an unsupported type.
-
ValueError
–
If either file contains invalid JSON, the lists have different
lengths, an entry has missing or extra keys, or paired filenames
differ.
|
Notes
Pass constructor arguments by keyword. len(benchmark) gives the number
of paired images. Construction loads the saved labels into memory and
leaves both source files unchanged. Subsequent edits to the source files
do not update an existing benchmark object.
BinaryClassificationBenchmark
dataclass
BinaryClassificationBenchmark(*, true_path: Path, pred_path: Path)
Bases: ClassificationBenchmark[int]
Evaluate binary image classification using saved labels 0 and 1.
The benchmark treats 1 as positive and 0 as negative. Construction
loads and validates the saved labels; metric properties compare those
labels without opening images or making model requests.
| Parameters: |
-
true_path
(Path)
–
Ground-truth JSON file containing a list of objects with exactly
img_path and label keys. Each img_path must be a string and each
label must be the integer 0 or 1. Boolean labels are rejected.
-
pred_path
(Path)
–
Prediction JSON file in the same format as true_path, such as
classifications.json saved by
classify_imgs(). Both lists must
contain the same number of entries in the same image order. At each
position, filenames including extensions must match; parent
directories may differ.
|
| Attributes: |
-
tp
(tuple[SingleClassificationResult[int], ...])
–
True-positive records: ground-truth label 1, predicted label 1.
-
fp
(tuple[SingleClassificationResult[int], ...])
–
False-positive records: ground-truth label 0, predicted label 1.
-
fn
(tuple[SingleClassificationResult[int], ...])
–
False-negative records: ground-truth label 1, predicted label 0.
-
tn
(tuple[SingleClassificationResult[int], ...])
–
True-negative records: ground-truth label 0, predicted label 0.
-
num_tp
(int)
–
Number of true-positive records.
-
num_fp
(int)
–
Number of false-positive records.
-
num_fn
(int)
–
Number of false-negative records.
-
num_tn
(int)
–
Number of true-negative records.
-
precision
(float)
–
Proportion of predicted positives that are correct:
num_tp / (num_tp + num_fp).
-
recall
(float)
–
Proportion of ground-truth positives found:
num_tp / (num_tp + num_fn).
-
f1_score
(float)
–
Harmonic mean of precision and recall:
2 * precision * recall / (precision + recall).
-
tpr
(float)
–
True-positive rate; equal to recall.
-
fpr
(float)
–
Proportion of ground-truth negatives incorrectly predicted positive:
num_fp / (num_fp + num_tn).
-
specificity
(float)
–
Proportion of ground-truth negatives correctly predicted negative:
num_tn / (num_tn + num_fp).
-
tnr
(float)
–
True-negative rate; equal to specificity.
|
| Raises: |
-
OSError
–
If either JSON file cannot be opened.
-
TypeError
–
If a JSON document is not a list, an entry is not an object, an
img_path is not a string, or a label is not an integer.
-
ValueError
–
If either file contains invalid JSON, the lists have different
lengths, an entry has missing or extra keys, paired filenames differ,
or a label is an integer other than 0 or 1.
|
Notes
Pass constructor arguments by keyword. All metric properties return
0.0 when their denominator is zero, including for two empty JSON lists.
Result tuples retain the JSON list order. Neither input file is changed.
The benchmark also exposes accuracy, result_list, img_paths,
trues, preds, correct_predictions and incorrect_predictions;
len(benchmark) counts paired images. See the shared properties on
ClassificationBenchmark.
Examples:
Using existing ground-truth and prediction JSON files:
>>> from pathlib import Path
>>> from flowde.benchmarks.classification.classification_benchmark import (
... BinaryClassificationBenchmark,
... )
>>> benchmark = BinaryClassificationBenchmark(
... true_path=Path("data/true-classifications.json"),
... pred_path=Path("results/classification/classifications.json"),
... )
>>> print(benchmark.accuracy, benchmark.precision, benchmark.recall)
MulticlassClassificationBenchmark
dataclass
MulticlassClassificationBenchmark(*, true_path: Path, pred_path: Path, labels: tuple[LabelType, ...])
Bases: ClassificationBenchmark[LabelType]
Evaluate image classification across a declared set of class labels.
The benchmark compares saved predicted labels with ground-truth labels
and provides overall accuracy, a confusion matrix and per-class counts.
Construction reads JSON files without opening images or making model
requests.
| Parameters: |
-
true_path
(Path)
–
Ground-truth JSON file containing a list of objects with exactly
img_path and label keys. Image paths must be strings; labels must
be str, int or bool values included in labels.
-
pred_path
(Path)
–
Prediction JSON file in the same format as true_path, such as
classifications.json saved by
classify_imgs(). Both lists must
contain the same number of entries in the same image order. At each
position, filenames including extensions must match; parent
directories may differ.
-
labels
(tuple[LabelType, ...])
–
Non-empty tuple of unique class labels. Must include every label in
both JSON files; may also include classes absent from both files.
Tuple order determines the order of dictionary keys in the confusion
matrix and class summaries. Uniqueness follows Python equality, so
1 and True, for example, cannot be separate classes.
|
| Attributes: |
-
confusion_matrix
(dict[LabelType, dict[LabelType, int]])
–
Counts indexed first by ground-truth label, then by predicted label.
confusion_matrix["table"]["flowchart"] counts true tables predicted
as flowcharts. Includes every pair of declared labels, with zero for
pairs absent from the saved results.
-
num_per_true_class
(dict[LabelType, int])
–
Number of images with each ground-truth label, including zero counts.
-
num_per_pred_class
(dict[LabelType, int])
–
Number of images with each predicted label, including zero counts.
-
per_class_accuracy
(dict[LabelType, float])
–
For each ground-truth class, the number of correct predictions divided
by the number of images in that class. Returns 0.0 for a declared
class with no ground-truth images.
|
| Raises: |
-
OSError
–
If either JSON file cannot be opened.
-
TypeError
–
If a JSON document is not a list, an entry is not an object, an
img_path is not a string, or a saved label has an unsupported type.
-
ValueError
–
If either file contains invalid JSON, the lists have different
lengths, an entry has missing or extra keys, paired filenames differ,
labels is empty or contains duplicates, or a saved label is absent
from labels.
|
Notes
Pass constructor arguments by keyword. Result tuples retain the JSON
list order. Neither input file is changed. Two empty JSON lists are
accepted with a non-empty labels tuple and produce zero counts and
accuracies.
The benchmark also exposes accuracy, result_list, img_paths,
trues, preds, correct_predictions and incorrect_predictions;
len(benchmark) counts paired images. See the shared properties on
ClassificationBenchmark.
Examples:
Using existing ground-truth and prediction JSON files:
>>> from pathlib import Path
>>> from flowde.benchmarks.classification.classification_benchmark import (
... MulticlassClassificationBenchmark,
... )
>>> benchmark = MulticlassClassificationBenchmark(
... true_path=Path("data/true-classifications.json"),
... pred_path=Path("results/classification/classifications.json"),
... labels=("flowchart", "table", "other"),
... )
>>> print(benchmark.confusion_matrix)
>>> print(benchmark.per_class_accuracy)
RotationBenchmark
RotationBenchmark(*, true_path: Path, pred_path: Path)
Bases: MulticlassClassificationBenchmark[RotationLabel]
Compare predicted image rotations with ground-truth rotation labels.
Labels describe the clockwise correction, in degrees, needed by each
original input image: 0, 90, 180 or 270. A prediction is correct
only when the predicted angle equals the ground-truth angle. The benchmark
reads saved labels without opening or rotating images or making model
requests.
| Parameters: |
-
true_path
(Path)
–
Ground-truth JSON file containing a list of objects with exactly
img_path and label keys. Each img_path must be a string; each
label specifies the clockwise correction for the original image.
-
pred_path
(Path)
–
Prediction JSON file in the same format as true_path, such as
rotations.json saved by
rotate_imgs(). Both lists must
contain the same number of entries in the same image order. At each
position, filenames including extensions must match; parent
directories may differ.
|
| Attributes: |
-
labels
(tuple[RotationLabel, ...])
–
Fixed class labels (0, 90, 180, 270). The constructor does not accept
a labels argument.
-
confusion_matrix
(dict[RotationLabel, dict[RotationLabel, int]])
–
Counts indexed first by ground-truth angle, then by predicted angle.
For example, confusion_matrix[90][0] counts images requiring a
90-degree correction that received a prediction of zero degrees.
Includes all four angles, even when no image has a particular angle.
-
num_per_true_class
(dict[RotationLabel, int])
–
Number of images requiring each ground-truth correction.
-
num_per_pred_class
(dict[RotationLabel, int])
–
Number of images assigned each predicted correction.
-
per_class_accuracy
(dict[RotationLabel, float])
–
For each ground-truth angle, correct predictions divided by images
requiring that angle. Returns 0.0 when no image requires the angle.
|
| Raises: |
-
OSError
–
If either JSON file cannot be opened.
-
TypeError
–
If a JSON document is not a list, an entry is not an object, an
img_path is not a string, or a saved label has an unsupported type.
-
ValueError
–
If either file contains invalid JSON, the lists have different
lengths, an entry has missing or extra keys, paired filenames differ,
or a saved label is outside the supported rotation classes.
|
Notes
Pass true_path and pred_path by keyword. Result tuples retain the JSON
list order. Neither input file is changed. Two empty JSON lists produce
zero counts and accuracies. Incorrect angles receive no partial credit
based on how close the predicted angle is to the ground-truth angle.
The benchmark also exposes accuracy, result_list, img_paths,
trues, preds, correct_predictions and incorrect_predictions;
len(benchmark) counts paired images. See the shared properties on
ClassificationBenchmark.
Examples:
Using existing ground-truth and prediction JSON files:
>>> from pathlib import Path
>>> from flowde.benchmarks.rotation.rotation_benchmark import RotationBenchmark
>>> benchmark = RotationBenchmark(
... true_path=Path("data/true-rotations.json"),
... pred_path=Path("results/rotation/rotations.json"),
... )
>>> print(benchmark.accuracy)
>>> print(benchmark.confusion_matrix)
Individual classification results
SingleClassificationResult
dataclass
SingleClassificationResult(*, img_path: str, pred: LabelType, true: LabelType)
Bases: Generic[LabelType]
Store the true and predicted labels for one benchmark image.
| Parameters: |
-
img_path
(str)
–
Path to the classified image. Benchmarks take this path from the
ground-truth classification file.
-
pred
(LabelType)
–
Predicted label for the image.
-
true
(LabelType)
–
Ground-truth label for the image.
|
Parsing
ParsingBenchmark
ParsingBenchmark(pred_diagrams_dir: Path | Sequence[Path], distance_fn: DistanceFnProtocol = levenshtein_with_text_normalisation, allow_missing_pred_diagrams: bool = False, true_nodes_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-nodes-texts-numbers-row-major-column/'), true_labels_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-labels-attempt-2/'), true_additional_texts_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-additional-texts/'), true_flow_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-flow/'), expected_num_diagrams: int = N_DIAGRAMS_IN_BENCHMARK)
Evaluate parsed flowcharts against accepted ground-truth interpretations.
Construction loads and validates saved JSON files. Scoring methods match
predicted nodes to ground-truth nodes and compare node text, labels,
directed connections and additional text. Benchmarking runs locally,
without opening images, changing source files or making model requests.
| Parameters: |
-
pred_diagrams_dir
(Path | Sequence[Path])
–
One directory containing prediction JSONs, or a sequence of directories
containing separately parsed parts. Reads top-level *.json files;
subdirectories are not searched. Each filename stem identifies a
flowchart, and each file contains one prediction object rather than
a ground-truth options list.
Predictions must include node text. Labels, flow and additional text
are optional. All JSONs within a prediction directory must contain
the same parts. Separate prediction directories must contain matching
filename stems and non-overlapping parts. Node-based parts must use
the same node numbers for each flowchart. The benchmark joins those
parts in memory; combined JSON files are not required.
-
distance_fn
(DistanceFnProtocol, default:
levenshtein_with_text_normalisation
)
–
Function accepting true_text and pred_text as keyword arguments
and returning a finite, non-negative int or float. Either argument
may be None to represent unmatched text; the supplied functions
treat None as an empty string. Lower costs mean closer text matches.
Defaults to
levenshtein_with_text_normalisation(),
which normalises Unicode, whitespace and selected punctuation before
counting character edits. You can supply a different function to
change the text comparison used for matching and scoring.
-
allow_missing_pred_diagrams
(bool, default:
False
)
–
Defaults to False, requiring predictions for every ground-truth
flowchart. With True, evaluate only flowcharts that have predictions.
Missing predictions receive no penalty and do not contribute to
averages. Predictions without corresponding ground truth remain
invalid.
-
true_nodes_dir
(Path, default:
Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-nodes-texts-numbers-row-major-column/')
)
–
Directory of ground-truth node-text JSONs. Each file contains an
options list of objects with nodes; each node supplies
node_number and text.
-
true_labels_dir
(Path, default:
Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-labels-attempt-2/')
)
–
Directory of ground-truth label JSONs. Each file contains an options
list of objects with nodes; each node supplies node_number and a
labels list of strings, which may be empty.
-
true_additional_texts_dir
(Path, default:
Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-additional-texts/')
)
–
Directory of ground-truth additional-text JSONs. Each file contains
an options list of objects with an additional_texts list of
strings, which may be empty.
-
true_flow_dir
(Path, default:
Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-flow/')
)
–
Directory of ground-truth flow JSONs. Each file contains an options
list of objects with nodes; each node supplies node_number and a
points_to list of destination node numbers.
-
expected_num_diagrams
(int, default:
N_DIAGRAMS_IN_BENCHMARK
)
–
Number of flowcharts in the complete ground-truth dataset, before
filtering missing predictions. Counts filename stems, not accepted
interpretations. Defaults to 346, the original research dataset's
size. Supply the count for your own dataset.
|
| Attributes: |
-
pred_diagrams
(list[Diagram])
–
Loaded predictions, including any parts joined from separate
directories, in sorted filename-stem order.
-
true_diagrams_options_list
(list[DiagramOptions])
–
Accepted ground-truth interpretations for each evaluated flowchart,
in the same filename-stem order as pred_diagrams.
-
pred_structure
(PredDiagramStructure)
–
Which parts the predictions contain: node text, labels, flow and
additional text.
-
capabilities
(ParsingBenchmarkCapabilities)
–
Which scoring operations the predicted parts support. The
available_methods tuple lists the supported method names.
|
| Raises: |
-
OSError
–
If a discovered JSON file cannot be opened.
-
TypeError
–
If a prediction or ground-truth field has an invalid type.
-
ValueError
–
If JSON is malformed, prediction directories are empty, file counts
or stems disagree, Ground-Truth Options do not align, or a node or
connection fails validation. Also raised for missing predicted node
text, overlapping predicted parts, or inconsistent predicted parts
or node numbers across their source files.
|
Notes
Supply all four ground-truth directories for your dataset, even when the
predictions contain only some parts. The default ground-truth paths refer
to the original research dataset, which is not included in a package
installation. The four directories must contain matching filename stems.
For each flowchart, the four options lists must have the same length;
entries at the same index form one Ground-Truth Option and must agree on
node numbers across node-based parts.
Scoring methods select Node Matches by minimising node-text cost. Among
ties, the benchmark maximises flow similarity, then uses label cost to
resolve remaining ties. Missing predicted parts are skipped. Ground-Truth
Option selection follows the same order, then uses additional-text cost
to resolve remaining ties; a complete tie selects the first option.
All scoring methods use the available predicted parts for matching.
Prediction and ground-truth node numbers do not need to match.
Text-cost methods sum matched and unmatched text costs. Flow-score methods
compare directed connections after matching nodes. Results follow sorted
filename-stem order, and avg_flow_jaccard() gives each evaluated
flowchart equal weight. Methods requiring an unparsed part raise
ValueError. A scoring method raises RuntimeError if the node matcher
cannot establish an optimal solution; an invalid text-distance value
raises ValueError during scoring.
Examples:
Using existing predictions and a ground-truth dataset of one flowchart:
>>> from pathlib import Path
>>> from flowde.benchmarks.parsing.parsing_bench import ParsingBenchmark
>>> truth_dir = Path("data/ground-truth")
>>> benchmark = ParsingBenchmark(
... pred_diagrams_dir=Path("results/predictions"),
... true_nodes_dir=truth_dir / "nodes",
... true_labels_dir=truth_dir / "labels",
... true_flow_dir=truth_dir / "flow",
... true_additional_texts_dir=truth_dir / "additional_texts",
... expected_num_diagrams=1,
... )
>>> print(benchmark.total_node_text_cost())
>>> print(benchmark.avg_flow_jaccard())
Methods
Select a method below for its parameters, return values and error conditions.
Select a result type to inspect the fields and properties on each returned
object. Lists follow sorted flowchart filename-stem order.
All methods use the available predicted parts to select Node Matches and a
Ground-Truth Option; see how diagrams are matched.
Each method calculates its results from the loaded data when called. Results
are not cached between calls, and scoring makes no model requests.
node_matches
node_matches(range_indices: tuple[int | None, int | None] | None = None) -> list[NodeMatches]
Match predicted nodes to ground-truth nodes for each flowchart.
Requires predicted node text. Matching minimises node-text cost,
uses flow similarity to resolve ties when flow was parsed, and then
uses label cost when labels were parsed. Ground-Truth Option selection
also uses additional-text cost to resolve remaining ties when that
part was parsed. A complete tie selects the first Ground-Truth Option.
| Parameters: |
-
range_indices
(tuple[int | None, int | None] | None, default:
None
)
–
A (start, stop) slice of the evaluated flowcharts in sorted
filename-stem order. start is included and stop is excluded.
Either bound may be None; negative indices follow Python
slicing rules. Defaults to None, which matches all evaluated
flowcharts. An empty slice returns an empty list.
|
| Returns: |
-
list[NodeMatches]
–
One collection per selected flowchart, in sorted filename-stem
order. Each collection identifies the selected Ground-Truth
Option and records paired nodes, unmatched ground-truth nodes and
unmatched predicted nodes, together with their node-text costs.
Node numbers identify nodes within each diagram; matching does
not require equal predicted and ground-truth node numbers.
|
| Raises: |
-
ValueError
–
If predicted node text is missing or a text-distance value fails
validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
Notes
The method uses every available predicted part for matching, even
though the returned result focuses on nodes. Label matches may be
populated while resolving ties; use
node_and_label_matches()
when you need label matches for every node.
node_matches_with_flow
node_matches_with_flow() -> list[NodeMatches]
Return Node Matches after requiring predicted flow as well as text.
Uses the same matching policy as
node_matches().
The additional flow requirement ensures that each returned collection
can calculate its flow_score property.
| Returns: |
-
list[NodeMatches]
–
One collection per evaluated flowchart, in sorted filename-stem
order, including the selected Ground-Truth Option and unmatched
nodes. Flow compares directed connections after applying the
Node Matches.
|
| Raises: |
-
ValueError
–
If predicted node text or flow is missing, or a text-distance
value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
node_and_label_matches
node_and_label_matches() -> list[NodeMatches]
Match nodes, then match the labels belonging to each Node Match.
Requires predicted node text and labels. The benchmark selects nodes
and a Ground-Truth Option using all available predicted parts, then
pairs label strings within each Node Match to minimise label cost.
Labels belonging to unmatched nodes count as unmatched labels.
| Returns: |
-
list[NodeMatches]
–
One collection per evaluated flowchart, in sorted filename-stem
order. Every NodeMatch.label_matches contains a TextListMatches
object, including an empty collection when both nodes have no
labels. total_label_error_cost sums the collection's label costs.
|
| Raises: |
-
ValueError
–
If predicted node text or labels are missing, or a text-distance
value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
additional_text_matches
additional_text_matches() -> list[TextListMatches]
Match additional-text strings for each evaluated flowchart.
Requires predicted node text and additional text. The benchmark first
selects a Ground-Truth Option using all available predicted parts,
then pairs that option's additional text with the prediction to
minimise additional-text cost. Strings are matched by cost, not by
their positions in the input lists.
| Returns: |
-
list[TextListMatches]
–
One collection per flowchart, in sorted filename-stem order.
Each collection contains paired and unmatched strings with their
original list indices and costs. Individual records also identify
the flowchart and selected Ground-Truth Option. If both diagrams
have empty additional-text lists, the collection has no records
and total_cost is zero.
|
| Raises: |
-
ValueError
–
If predicted node text or additional text is missing, or a
text-distance value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
diagram_matches
diagram_matches() -> list[DiagramMatch]
Match all four parsed parts for each evaluated flowchart.
Requires predicted node text, labels, flow and additional text. The
benchmark selects one Ground-Truth Option per prediction and uses
that option for the node, label, additional-text and flow comparisons.
| Returns: |
-
list[DiagramMatch]
–
One complete comparison per flowchart, in sorted filename-stem
order. Each result contains the prediction, the selected
ground-truth diagram, Node Matches with label matches, and
additional-text matches. Text totals and flow scores are
available as properties on each result.
|
| Raises: |
-
ValueError
–
If any required predicted part is missing or a text-distance
value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
total_node_text_cost
total_node_text_cost() -> float
Sum node-text costs across all evaluated flowcharts.
| Returns: |
-
float
–
Total cost of paired and unmatched node text, using distance_fn.
With the default distance function, the total counts character
edits after text normalisation. Lower is better; the total is not
an average or percentage. Requires predicted node text.
|
| Raises: |
-
ValueError
–
If predicted node text is missing or a text-distance value fails
validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
total_label_cost
total_label_cost() -> float
Sum label costs across all evaluated flowcharts.
| Returns: |
-
float
–
Total cost of paired and unmatched label strings, including
labels on unmatched nodes, using distance_fn. Lower is better;
the total is not an average or percentage. Requires predicted
node text and labels. Node text selects the Node Matches but
does not contribute to this label-only total.
|
| Raises: |
-
ValueError
–
If predicted node text or labels are missing, or a text-distance
value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
total_additional_text_cost
total_additional_text_cost() -> float
Sum additional-text costs across all evaluated flowcharts.
| Returns: |
-
float
–
Total cost of paired and unmatched additional-text strings, using
distance_fn. Lower is better; the total is not an average or
percentage. Requires predicted node text and additional text.
Node text helps select the Ground-Truth Option but does not
contribute to this additional-text-only total.
|
| Raises: |
-
ValueError
–
If predicted node text or additional text is missing, or a
text-distance value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
all_flow_scores
all_flow_scores() -> list[FlowScores]
Score the predicted directed connections in each flowchart.
Requires predicted node text and flow. The benchmark matches nodes
before comparing connections, so predicted node numbers may differ
from ground-truth node numbers. Connection direction matters:
1 -> 2 and 2 -> 1 represent different connections.
| Returns: |
-
list[FlowScores]
–
One score object per evaluated flowchart, in sorted filename-stem
order. Each object contains correct, extra and missing connection
counts, precision, recall, F1, Jaccard similarity, and the sets of
extra and missing connections. Reported connections use matched
ground-truth node numbers; unmatched predicted endpoints receive
generated identifiers starting at 10000.
|
| Raises: |
-
ValueError
–
If predicted node text or flow is missing, or a text-distance
value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
all_flow_jaccard_scores
all_flow_jaccard_scores() -> list[float]
Return one flow Jaccard similarity score per evaluated flowchart.
| Returns: |
-
list[float]
–
Scores in sorted filename-stem order, calculated after matching
nodes. Each score is TP / (TP + FP + FN), between 0.0 and
1.0; higher is better. A score of 1.0 means the directed
connection sets agree. Two empty connection sets score 1.0.
Requires predicted node text and flow.
|
| Raises: |
-
ValueError
–
If predicted node text or flow is missing, or a text-distance
value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
avg_flow_jaccard
avg_flow_jaccard() -> float
Average the flow Jaccard scores of the evaluated flowcharts.
| Returns: |
-
float
–
Arithmetic mean of the individual flow Jaccard scores, between
0.0 and 1.0; higher is better. Each flowchart has equal weight,
regardless of its number of connections. Missing predictions
excluded with allow_missing_pred_diagrams=True do not contribute
to the mean. Requires predicted node text and flow.
|
| Raises: |
-
ValueError
–
If predicted node text or flow is missing, there are no evaluated
flowcharts, or a text-distance value fails validation.
-
RuntimeError
–
If the node matcher cannot establish an optimal solution.
|
Matching result objects
These Pydantic models hold the comparisons returned by the benchmark.
A match can pair a prediction with ground truth or record an unmatched node
or string. The following entries describe the stored fields and computed
properties.
| Type |
Contains |
NodeMatches |
Node Matches for one flowchart and the selected Ground-Truth Option. |
NodeMatch |
One paired or unmatched node, with text cost and optional label matches. |
TextListMatches |
Comparisons for one node's labels or one flowchart's additional text. |
TextListMatch |
One paired or unmatched string, with source-list indices and cost. |
DiagramMatch |
A complete prediction, its selected ground-truth diagram and all part comparisons. |
FlowScores |
Directed-connection counts, metrics, and missing and extra connections. |
Node |
An individual predicted or ground-truth node held inside a match. |
Diagram |
The prediction or selected ground-truth diagram held inside a complete comparison. |
You can inspect stored fields with
model_dump()
or serialise stored fields with
model_dump_json().
Computed properties such as total_text_cost and flow_score are accessed
directly and are not included in those dumps.
NodeMatches
Bases: BaseModel
Collect the Node Matches for one flowchart and one Ground-Truth Option.
Benchmark-generated collections include every ground-truth and predicted
node exactly once, either in a pair or as an unmatched node.
| Attributes: |
-
matches
(list[NodeMatch])
–
Paired and unmatched node records. Inspect true_node and pred_node
on each record to identify the nodes; list positions are not node
numbers.
-
true_diagram_option_idx
(int)
–
Index of the selected Ground-Truth Option, starting at 0.
-
parent_img_code
(str)
–
Source flowchart's filename stem, taken from the first Node Match.
Access raises ValueError if matches is empty.
-
true_nodes
(list[Node])
–
All present ground-truth nodes, including unmatched nodes, in the
order of their records in matches.
-
pred_nodes
(list[Node])
–
All present predicted nodes, including unmatched nodes, in the
order of their records in matches.
-
total_node_text_cost
(float)
–
Sum of the records' node_text_cost values, including unmatched nodes.
-
total_numbers_only_node_text_cost
(float)
–
Sum of the records' numbers-only text costs. Uses the existing Node
Matches without choosing different pairings or a different option.
-
total_label_error_cost
(float)
–
Sum of label costs across all Node Matches. Access raises ValueError
if any record has label_matches=None; an empty collection costs zero.
-
flow_score
(FlowScores)
–
Directed-connection scores calculated from these Node Matches.
Requires a points_to list on every present node. Matched endpoints
use ground-truth node numbers; unmatched predicted endpoints receive
generated identifiers starting at 10000.
|
| Raises: |
-
ValidationError
–
If a field is invalid or a ground-truth or predicted node number
appears in more than one record on the same side during validation.
|
Notes
Text costs include unmatched nodes and labels. Flow uses the same Node
Matches; accessing flow_score does not select new pairings. Cost
properties sum the stored records rather than rerunning model requests
or matching.
get_matched_true_node
get_matched_true_node(pred_node_number: int, allow_fake_pred_node: bool = False) -> Node | None
Find the ground-truth node paired with a predicted node number.
| Parameters: |
-
pred_node_number
(int)
–
Identifier of a predicted node recorded in matches.
-
allow_fake_pred_node
(bool, default:
False
)
–
Whether an absent predicted node number should return None
instead of raising an error. Defaults to False. Flow scoring
uses this option when looking up predicted connection endpoints.
|
| Returns: |
-
Node | None
–
The paired ground-truth node. Returns None for a recorded
unmatched predicted node, or for an absent predicted node number
when allow_fake_pred_node=True.
|
| Raises: |
-
ValueError
–
If pred_node_number is absent and allow_fake_pred_node=False.
|
get_matched_pred_node
get_matched_pred_node(true_node_number: int) -> Node | None
Find the predicted node paired with a ground-truth node number.
| Parameters: |
-
true_node_number
(int)
–
Identifier of a ground-truth node recorded in matches.
|
| Returns: |
-
Node | None
–
The paired predicted node, or None when the recorded ground-truth
node is unmatched.
|
| Raises: |
-
ValueError
–
If true_node_number does not appear in matches.
|
NodeMatch
Bases: BaseModel
Record one pair of corresponding nodes or one unmatched node.
A paired prediction and ground-truth node may have different text or node
numbers. The benchmark stores text cost separately from any label costs.
| Attributes: |
-
true_node
(Node | None)
–
Node from the selected Ground-Truth Option, or None when a predicted
node has no ground-truth match.
-
pred_node
(Node | None)
–
Predicted node, or None when a ground-truth node has no prediction.
-
node_text_cost
(int | float)
–
Cost returned by the benchmark's distance_fn for the node text.
An unmatched node is compared with a missing text value; built-in
distance functions treat the missing value as an empty string.
Excludes label and flow costs.
-
label_matches
(TextListMatches | None)
–
Label comparisons for this Node Match, or None before labels have
been matched. An empty collection means both nodes have no labels.
Label comparisons include labels belonging to unmatched nodes.
-
match_type
({'match', 'unmatched_true', 'unmatched_pred'})
–
"match" for two present nodes; "unmatched_true" for a ground-truth
node without a prediction; "unmatched_pred" for a predicted node
without ground truth. A pair need not have zero text cost.
-
parent_img_code
(str)
–
Source flowchart's filename stem, taken from the ground-truth node
when present, otherwise from the predicted node.
-
numbers_only_node_text_cost
(float)
–
Cost from comparing only the numbers in the two nodes' text with
number_only_levenshtein().
Uses this existing Node Match without changing the nodes or their
pairing.
|
| Raises: |
-
ValidationError
–
If a field is invalid, both nodes are absent, or two present nodes
have different parent_img_code values.
|
TextListMatches
Bases: BaseModel
Collect comparisons for one node's labels or a diagram's additional text.
The benchmark matches strings to minimise total cost rather than pairing
strings by their list positions. Each record retains the original indices
so the source strings can be identified.
| Attributes: |
-
matches
(list[TextListMatch])
–
Paired and unmatched string records. Benchmark-generated collections
account for every string in both source lists. Two empty source lists
produce an empty collection. Use each record's true_index and
pred_index to identify its original positions.
-
total_cost
(float)
–
Sum of every record's cost, including unmatched strings. An empty
collection has cost zero. The property reports a total, not an average.
|
Notes
This class stores the supplied records and computes their total; the
class does not perform string matching when constructed. The container
does not store source lists or flowchart metadata. For additional text,
the individual records hold the flowchart and Ground-Truth Option IDs.
TextListMatch
Bases: BaseModel
Record one paired or unmatched label or additional-text string.
A pair records which predicted string represents which ground-truth
string, even when their text differs. An unmatched record contains text
on only one side.
| Attributes: |
-
true_index
(int | None)
–
Zero-based position in the ground-truth text list. None when a
predicted string has no ground-truth match.
-
pred_index
(int | None)
–
Zero-based position in the predicted text list. None when a
ground-truth string has no predicted match.
-
true_text
(str | None)
–
Original ground-truth string, before normalisation; None exactly
when true_index is None. Present strings must be non-empty.
-
pred_text
(str | None)
–
Original predicted string, before normalisation; None exactly when
pred_index is None. Present strings must be non-empty.
-
parent_img_code
(str | None)
–
Source flowchart's filename stem. Additional-text matching sets this
field; label matching leaves the default None because the containing
NodeMatch identifies the flowchart.
-
true_option_idx
(int | None)
–
Selected Ground-Truth Option index, starting at 0. Additional-text
matching sets this field; label matching leaves the default None.
-
cost
(float)
–
Distance-function cost for this string pair or unmatched string.
An unmatched string is compared with None; the built-in distance
functions treat the missing side as an empty string.
-
match_type
({'match', 'unmatched_true', 'unmatched_pred'})
–
"match" when both strings are present; "unmatched_true" for a
ground-truth string without a prediction; "unmatched_pred" for a
predicted string without ground truth. "match" does not imply
identical text or zero cost.
|
| Raises: |
-
ValidationError
–
If a field is invalid, an index and its text disagree about whether
the corresponding side is absent, or both sides are absent.
|
DiagramMatch
Bases: BaseModel
Collect a complete prediction's comparison with one Ground-Truth Option.
Both diagrams contain node text, labels, flow and additional text. The
result exposes the loaded diagrams, their matching records, and the text
and flow scores calculated from those records.
| Attributes: |
-
node_matches
(NodeMatches)
–
Every paired and unmatched node, with label matches for each record.
true_diagram_option_idx identifies the selected Ground-Truth Option.
-
additional_text_matches
(TextListMatches)
–
Paired and unmatched additional-text strings from both diagrams.
-
pred_diagram
(Diagram)
–
Complete predicted diagram, retaining its original node numbers and
text.
-
true_diagram
(Diagram)
–
Complete ground-truth diagram selected from the accepted options.
true_option_idx is its zero-based option index.
-
total_node_text_cost
(float)
–
Sum of paired and unmatched node-text costs.
-
total_label_error_cost
(float)
–
Sum of paired and unmatched label costs, including labels on unmatched
nodes.
-
total_additional_text_cost
(float)
–
Sum of paired and unmatched additional-text costs.
-
total_text_cost
(float)
–
Sum of node-text, label and additional-text costs. Excludes flow scores.
-
flow_score
(FlowScores)
–
Directed-connection scores after applying node_matches.
|
| Raises: |
-
ValidationError
–
If a field is invalid, either diagram is incomplete, the diagrams
refer to different flowcharts or have incorrect diagram_type
values, or the node and additional-text records do not account for
the contents of the compared diagrams.
|
Flow scores
FlowScores
Bases: BaseModel
Store directed-connection counts and scores for one matched flowchart.
The benchmark calculates these fields after matching predicted nodes to
ground-truth nodes. Connection direction matters: 1 -> 2 differs from
2 -> 1.
| Attributes: |
-
tp
(int)
–
Number of connections present in both the prediction and ground truth.
-
fp
(int)
–
Number of predicted connections absent from ground truth.
-
fn
(int)
–
Number of ground-truth connections absent from the prediction.
-
precision
(float)
–
Correct connections divided by predicted connections:
tp / (tp + fp). Zero when the prediction has no connections.
-
recall
(float)
–
Correct connections divided by ground-truth connections:
tp / (tp + fn). Zero when ground truth has no connections.
-
f1
(float)
–
Harmonic mean of precision and recall. Zero when both are zero.
-
jaccard
(float)
–
Correct connections divided by all distinct connections in either
diagram: tp / (tp + fp + fn). Ranges from 0.0 to 1.0; higher is
better. Two empty connection sets score 1.0.
-
missing_edges
(set[tuple[int, int]])
–
Missing ground-truth connections as (source, destination) pairs
of ground-truth node numbers.
-
extra_edges
(set[tuple[int, int]])
–
Predicted connections absent from ground truth, expressed using
matched ground-truth node numbers. Unmatched predicted endpoints use
generated identifiers starting at 10000, not original node numbers.
|
Notes
A FlowScores object stores supplied values; constructing the object does
not calculate metrics from tp, fp and fn. Benchmark methods calculate
the metrics before constructing the object.
Nodes and diagrams
These benchmark models preserve the source text and node numbers alongside
the metadata identifying the flowchart and Ground-Truth Option.
Node
Bases: BaseModel
Represent one predicted or ground-truth node inside a parsing benchmark.
Benchmark loaders attach the flowchart identifier and Ground-Truth Option
index to the node fields read from JSON. Node numbers identify nodes within
a diagram; predicted and ground-truth versions of a node may have different
numbers.
| Attributes: |
-
node_number
(int)
–
Node identifier. Must be less than 1000; the containing Diagram
checks uniqueness and applies the ground-truth numbering rules.
-
text
(str)
–
Non-empty node text as loaded, before any distance-function
normalisation.
-
labels
(list[str] | None)
–
Non-empty label strings. [] means labels were parsed and none were
found; None, the default, means labels were not parsed.
-
points_to
(list[int] | None)
–
Destination node numbers for outgoing connections. [] means flow
was parsed and the node has no outgoing connections; None, the
default, means flow was not parsed.
-
diagram_type
({'pred', 'true'})
–
Whether the node belongs to a prediction or to ground truth.
-
true_option_idx
(int | None)
–
Zero-based Ground-Truth Option index. Required for a ground-truth
node; must be None for a predicted node. Defaults to None.
-
parent_img_code
(str)
–
Filename stem identifying the source flowchart.
-
text_done
(bool)
–
Whether node text is present; always True for a valid Node.
-
labels_done
(bool)
–
Whether labels is present, including an empty list.
-
flow_done
(bool)
–
Whether points_to is present, including an empty list.
-
is_complete
(bool)
–
Whether text, labels and flow are all present.
|
| Raises: |
-
ValidationError
–
If a field has an invalid type, text or a label string is empty,
node_number is at least 1000, or true_option_idx conflicts with
diagram_type.
|
Diagram
Bases: BaseModel
Represent one prediction or one accepted ground-truth interpretation.
A DiagramMatch exposes the compared diagrams through pred_diagram and
true_diagram. Diagram fields retain the original text and node numbers;
matching and text normalisation do not rewrite the loaded diagrams.
| Attributes: |
-
nodes
(list[Node] | None)
–
Non-empty list of nodes, or None when no node-based parts were
supplied. Defaults to None. Parsing benchmarks require predicted
node text, so diagrams returned by benchmark methods contain nodes.
-
additional_texts
(list[str] | None)
–
Non-empty strings outside the nodes. [] means additional text was
parsed and none was found; None, the default, means the part was
not parsed.
-
parent_img_code
(str)
–
Filename stem identifying the source flowchart.
-
diagram_type
({'pred', 'true'})
–
Whether the diagram is a prediction or a Ground-Truth Option.
-
true_option_idx
(int | None)
–
Zero-based Ground-Truth Option index for a ground-truth diagram;
None for a predicted diagram.
-
text_done
(bool)
–
Whether nodes and their text are present.
-
labels_done
(bool)
–
Whether every node has a labels list, including empty lists.
-
flow_done
(bool)
–
Whether every node has a points_to list, including empty lists.
-
additional_texts_done
(bool)
–
Whether additional_texts is present, including an empty list.
-
is_complete
(bool)
–
Whether node text, labels, flow and additional text are all present.
|
| Raises: |
-
ValidationError
–
If fields are invalid, node numbers repeat, connections refer to
absent nodes, or node metadata disagrees with the diagram. Also
raised if only some nodes supply labels or flow, or a ground-truth
diagram is incomplete, has non-sequential node numbers, or contains
a node absent from every connection.
|
Notes
Ground-truth nodes must appear in consecutive node-number order starting
at 1. Predictions may use different numbering. Ground-truth diagrams
must contain all four parts; predictions may omit labels, flow or
additional text.
Text distances
Import the following functions from
flowde.benchmarks.parsing.text_distance_fns.levenshtein_fn.
levenshtein_with_text_normalisation
levenshtein_with_text_normalisation(true_text: str | None, pred_text: str | None) -> int
Count character edits after normalising both input strings.
This is the default text distance used by ParsingBenchmark. Normalisation
reduces differences caused by formatting and alternative character forms
before Levenshtein distance counts insertions, deletions and substitutions.
| Parameters: |
-
true_text
(str | None)
–
Ground-truth text. None represents unmatched predicted text and is
treated as an empty string.
-
pred_text
(str | None)
–
Predicted text. None represents missing predicted text and is treated
as an empty string.
|
| Returns: |
-
int
–
Minimum number of single-character edits between the normalised
strings. 0 means the normalised strings are identical. The result is
an edit count, not a score scaled between zero and one.
|
| Raises: |
-
ValueError
–
If both true_text and pred_text are None.
|
Notes
Both strings receive the same normalisation:
- Apply Unicode NFC so equivalent composed and decomposed characters match.
- Standardise line endings, convert tabs to spaces, collapse repeated
ordinary spaces and newlines, and strip surrounding whitespace.
- Standardise supported bullet, dash, quotation-mark and caret characters.
A line beginning with
o or O followed by whitespace becomes a bullet.
- Standardise spacing around punctuation, brackets, equals signs, plus
signs and hyphens between word characters, including digits.
- Collapse repeated HTML line-break tags to one
<br> tag.
- Convert superscript ordinal suffixes following digits, such as
1ˢᵗ
to 1st.
Normalisation does not lowercase the text or remove accents. Single
newlines remain distinct from spaces unless a punctuation-spacing rule
removes the newline. HTML break tags are not converted to newlines.
Examples:
>>> levenshtein_with_text_normalisation("n=10", "n = 10")
0
>>> levenshtein_with_text_normalisation("n = 10", "n = 11")
1
>>> levenshtein_with_text_normalisation(None, " ABC ")
3
ParsingBenchmark
uses this function when you omit distance_fn. You can still pass
levenshtein_fn()
for raw character comparisons, or another distance function.
levenshtein_fn
levenshtein_fn(true_text: str | None, pred_text: str | None) -> int
Counts the minimum number of single-character insertions, deletions and
substitutions. Treats a missing string as empty; two missing strings are invalid.
number_only_levenshtein
number_only_levenshtein(true_text: str | None, pred_text: str | None) -> int
Compares the extracted numbers in the two strings. Use this for a separate
number-focused check, not as a measure of whether the surrounding words agree.
NodeMatches.total_numbers_only_node_text_cost provides this check for the
already selected node matches.
DistanceFnProtocol
Bases: Protocol
Describe a callable that calculates a cost between true and predicted text.
Pass a function with this interface as distance_fn to
ParsingBenchmark
to compare node text, labels and additional text. Lower costs must indicate
closer matches because the benchmark minimises text costs when matching
predictions to ground truth.
Notes
This protocol describes the callable interface; it does not calculate
costs or validate returned values. A function or callable object can
satisfy the interface without inheriting from DistanceFnProtocol.
__call__
__call__(*, true_text: str | None, pred_text: str | None) -> int | float
Calculate the comparison cost for one pair of text values.
| Parameters: |
-
true_text
(str | None)
–
Ground-truth text. The benchmark supplies this argument by keyword.
None represents predicted text with no ground-truth match.
-
pred_text
(str | None)
–
Predicted text. The benchmark supplies this argument by keyword.
None represents ground-truth text with no prediction.
|
| Returns: |
-
int | float
–
Finite, non-negative comparison cost. Lower values must indicate
closer text matches. Return a cost rather than a similarity score
where higher values indicate closer matches.
|
Notes
The function must handle either argument being None and assign a
cost to the unmatched text. Flowde's built-in distance functions treat
None as an empty string. A custom function can choose another penalty
for unmatched text. The benchmark does not compare two None values.