Benchmarks

Benchmarks read saved JSONs and run locally. Constructing or evaluating a benchmark does not call an LLM. See the stage-specific benchmarking guides for input formats.

Classification and rotation

ClassificationBenchmark dataclass

ClassificationBenchmark(*, true_path: Path, pred_path: Path)

Bases: Generic[LabelType]

Compare saved image labels with ground truth and report accuracy.

The benchmark reads JSON files without opening images or making model requests. Binary, multiclass and rotation benchmarks inherit the shared result properties documented here.

Parameters:
  • true_path (Path) –

    Ground-truth JSON file containing a list of objects with exactly img_path and label keys. Image paths must be strings; labels must be str, int or bool values.

  • pred_path (Path) –

    Prediction JSON file with the same format and number of entries as true_path. Entries are paired by list position. Paired image filenames, including extensions, must match; parent directories may differ. The benchmark does not sort or reorder either list.

Attributes:
  • result_list (tuple[SingleClassificationResult[LabelType], ...]) –

    One record per paired image, in the JSON list order. Each record contains img_path from ground truth, true and pred.

  • correct_predictions (tuple[SingleClassificationResult[LabelType], ...]) –

    Records whose predicted labels equal their ground-truth labels.

  • incorrect_predictions (tuple[SingleClassificationResult[LabelType], ...]) –

    Records whose predicted labels differ from their ground-truth labels.

  • accuracy (float) –

    Number of correct predictions divided by the number of paired images. Returns 0.0 for two empty JSON lists.

  • img_paths (tuple[str, ...]) –

    Image paths from the ground-truth JSON, in the JSON list order.

  • trues (tuple[LabelType, ...]) –

    Ground-truth labels, in the JSON list order.

  • preds (tuple[LabelType, ...]) –

    Predicted labels, in the JSON list order.

Raises:
  • OSError –

    If either JSON file cannot be opened.

  • TypeError –

    If a JSON document is not a list, an entry is not an object, an img_path is not a string, or a label has an unsupported type.

  • ValueError –

    If either file contains invalid JSON, the lists have different lengths, an entry has missing or extra keys, or paired filenames differ.

Notes

Pass constructor arguments by keyword. len(benchmark) gives the number of paired images. Construction loads the saved labels into memory and leaves both source files unchanged. Subsequent edits to the source files do not update an existing benchmark object.

BinaryClassificationBenchmark dataclass

BinaryClassificationBenchmark(*, true_path: Path, pred_path: Path)

Bases: ClassificationBenchmark[int]

Evaluate binary image classification using saved labels 0 and 1.

The benchmark treats 1 as positive and 0 as negative. Construction loads and validates the saved labels; metric properties compare those labels without opening images or making model requests.

Parameters:
  • true_path (Path) –

    Ground-truth JSON file containing a list of objects with exactly img_path and label keys. Each img_path must be a string and each label must be the integer 0 or 1. Boolean labels are rejected.

  • pred_path (Path) –

    Prediction JSON file in the same format as true_path, such as classifications.json saved by classify_imgs(). Both lists must contain the same number of entries in the same image order. At each position, filenames including extensions must match; parent directories may differ.

Attributes:
  • tp (tuple[SingleClassificationResult[int], ...]) –

    True-positive records: ground-truth label 1, predicted label 1.

  • fp (tuple[SingleClassificationResult[int], ...]) –

    False-positive records: ground-truth label 0, predicted label 1.

  • fn (tuple[SingleClassificationResult[int], ...]) –

    False-negative records: ground-truth label 1, predicted label 0.

  • tn (tuple[SingleClassificationResult[int], ...]) –

    True-negative records: ground-truth label 0, predicted label 0.

  • num_tp (int) –

    Number of true-positive records.

  • num_fp (int) –

    Number of false-positive records.

  • num_fn (int) –

    Number of false-negative records.

  • num_tn (int) –

    Number of true-negative records.

  • precision (float) –

    Proportion of predicted positives that are correct: num_tp / (num_tp + num_fp).

  • recall (float) –

    Proportion of ground-truth positives found: num_tp / (num_tp + num_fn).

  • f1_score (float) –

    Harmonic mean of precision and recall: 2 * precision * recall / (precision + recall).

  • tpr (float) –

    True-positive rate; equal to recall.

  • fpr (float) –

    Proportion of ground-truth negatives incorrectly predicted positive: num_fp / (num_fp + num_tn).

  • specificity (float) –

    Proportion of ground-truth negatives correctly predicted negative: num_tn / (num_tn + num_fp).

  • tnr (float) –

    True-negative rate; equal to specificity.

Raises:
  • OSError –

    If either JSON file cannot be opened.

  • TypeError –

    If a JSON document is not a list, an entry is not an object, an img_path is not a string, or a label is not an integer.

  • ValueError –

    If either file contains invalid JSON, the lists have different lengths, an entry has missing or extra keys, paired filenames differ, or a label is an integer other than 0 or 1.

Notes

Pass constructor arguments by keyword. All metric properties return 0.0 when their denominator is zero, including for two empty JSON lists. Result tuples retain the JSON list order. Neither input file is changed.

The benchmark also exposes accuracy, result_list, img_paths, trues, preds, correct_predictions and incorrect_predictions; len(benchmark) counts paired images. See the shared properties on ClassificationBenchmark.

Examples:

Using existing ground-truth and prediction JSON files:

>>> from pathlib import Path
>>> from flowde.benchmarks.classification.classification_benchmark import (
...     BinaryClassificationBenchmark,
... )
>>> benchmark = BinaryClassificationBenchmark(
...     true_path=Path("data/true-classifications.json"),
...     pred_path=Path("results/classification/classifications.json"),
... )
>>> print(benchmark.accuracy, benchmark.precision, benchmark.recall)

MulticlassClassificationBenchmark dataclass

MulticlassClassificationBenchmark(*, true_path: Path, pred_path: Path, labels: tuple[LabelType, ...])

Bases: ClassificationBenchmark[LabelType]

Evaluate image classification across a declared set of class labels.

The benchmark compares saved predicted labels with ground-truth labels and provides overall accuracy, a confusion matrix and per-class counts. Construction reads JSON files without opening images or making model requests.

Parameters:
  • true_path (Path) –

    Ground-truth JSON file containing a list of objects with exactly img_path and label keys. Image paths must be strings; labels must be str, int or bool values included in labels.

  • pred_path (Path) –

    Prediction JSON file in the same format as true_path, such as classifications.json saved by classify_imgs(). Both lists must contain the same number of entries in the same image order. At each position, filenames including extensions must match; parent directories may differ.

  • labels (tuple[LabelType, ...]) –

    Non-empty tuple of unique class labels. Must include every label in both JSON files; may also include classes absent from both files. Tuple order determines the order of dictionary keys in the confusion matrix and class summaries. Uniqueness follows Python equality, so 1 and True, for example, cannot be separate classes.

Attributes:
  • confusion_matrix (dict[LabelType, dict[LabelType, int]]) –

    Counts indexed first by ground-truth label, then by predicted label. confusion_matrix["table"]["flowchart"] counts true tables predicted as flowcharts. Includes every pair of declared labels, with zero for pairs absent from the saved results.

  • num_per_true_class (dict[LabelType, int]) –

    Number of images with each ground-truth label, including zero counts.

  • num_per_pred_class (dict[LabelType, int]) –

    Number of images with each predicted label, including zero counts.

  • per_class_accuracy (dict[LabelType, float]) –

    For each ground-truth class, the number of correct predictions divided by the number of images in that class. Returns 0.0 for a declared class with no ground-truth images.

Raises:
  • OSError –

    If either JSON file cannot be opened.

  • TypeError –

    If a JSON document is not a list, an entry is not an object, an img_path is not a string, or a saved label has an unsupported type.

  • ValueError –

    If either file contains invalid JSON, the lists have different lengths, an entry has missing or extra keys, paired filenames differ, labels is empty or contains duplicates, or a saved label is absent from labels.

Notes

Pass constructor arguments by keyword. Result tuples retain the JSON list order. Neither input file is changed. Two empty JSON lists are accepted with a non-empty labels tuple and produce zero counts and accuracies.

The benchmark also exposes accuracy, result_list, img_paths, trues, preds, correct_predictions and incorrect_predictions; len(benchmark) counts paired images. See the shared properties on ClassificationBenchmark.

Examples:

Using existing ground-truth and prediction JSON files:

>>> from pathlib import Path
>>> from flowde.benchmarks.classification.classification_benchmark import (
...     MulticlassClassificationBenchmark,
... )
>>> benchmark = MulticlassClassificationBenchmark(
...     true_path=Path("data/true-classifications.json"),
...     pred_path=Path("results/classification/classifications.json"),
...     labels=("flowchart", "table", "other"),
... )
>>> print(benchmark.confusion_matrix)
>>> print(benchmark.per_class_accuracy)

RotationBenchmark

RotationBenchmark(*, true_path: Path, pred_path: Path)

Bases: MulticlassClassificationBenchmark[RotationLabel]

Compare predicted image rotations with ground-truth rotation labels.

Labels describe the clockwise correction, in degrees, needed by each original input image: 0, 90, 180 or 270. A prediction is correct only when the predicted angle equals the ground-truth angle. The benchmark reads saved labels without opening or rotating images or making model requests.

Parameters:
  • true_path (Path) –

    Ground-truth JSON file containing a list of objects with exactly img_path and label keys. Each img_path must be a string; each label specifies the clockwise correction for the original image.

  • pred_path (Path) –

    Prediction JSON file in the same format as true_path, such as rotations.json saved by rotate_imgs(). Both lists must contain the same number of entries in the same image order. At each position, filenames including extensions must match; parent directories may differ.

Attributes:
  • labels (tuple[RotationLabel, ...]) –

    Fixed class labels (0, 90, 180, 270). The constructor does not accept a labels argument.

  • confusion_matrix (dict[RotationLabel, dict[RotationLabel, int]]) –

    Counts indexed first by ground-truth angle, then by predicted angle. For example, confusion_matrix[90][0] counts images requiring a 90-degree correction that received a prediction of zero degrees. Includes all four angles, even when no image has a particular angle.

  • num_per_true_class (dict[RotationLabel, int]) –

    Number of images requiring each ground-truth correction.

  • num_per_pred_class (dict[RotationLabel, int]) –

    Number of images assigned each predicted correction.

  • per_class_accuracy (dict[RotationLabel, float]) –

    For each ground-truth angle, correct predictions divided by images requiring that angle. Returns 0.0 when no image requires the angle.

Raises:
  • OSError –

    If either JSON file cannot be opened.

  • TypeError –

    If a JSON document is not a list, an entry is not an object, an img_path is not a string, or a saved label has an unsupported type.

  • ValueError –

    If either file contains invalid JSON, the lists have different lengths, an entry has missing or extra keys, paired filenames differ, or a saved label is outside the supported rotation classes.

Notes

Pass true_path and pred_path by keyword. Result tuples retain the JSON list order. Neither input file is changed. Two empty JSON lists produce zero counts and accuracies. Incorrect angles receive no partial credit based on how close the predicted angle is to the ground-truth angle.

The benchmark also exposes accuracy, result_list, img_paths, trues, preds, correct_predictions and incorrect_predictions; len(benchmark) counts paired images. See the shared properties on ClassificationBenchmark.

Examples:

Using existing ground-truth and prediction JSON files:

>>> from pathlib import Path
>>> from flowde.benchmarks.rotation.rotation_benchmark import RotationBenchmark
>>> benchmark = RotationBenchmark(
...     true_path=Path("data/true-rotations.json"),
...     pred_path=Path("results/rotation/rotations.json"),
... )
>>> print(benchmark.accuracy)
>>> print(benchmark.confusion_matrix)

Individual classification results

SingleClassificationResult dataclass

SingleClassificationResult(*, img_path: str, pred: LabelType, true: LabelType)

Bases: Generic[LabelType]

Store the true and predicted labels for one benchmark image.

Parameters:
  • img_path (str) –

    Path to the classified image. Benchmarks take this path from the ground-truth classification file.

  • pred (LabelType) –

    Predicted label for the image.

  • true (LabelType) –

    Ground-truth label for the image.

Parsing

ParsingBenchmark

ParsingBenchmark(pred_diagrams_dir: Path | Sequence[Path], distance_fn: DistanceFnProtocol = levenshtein_with_text_normalisation, allow_missing_pred_diagrams: bool = False, true_nodes_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-nodes-texts-numbers-row-major-column/'), true_labels_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-labels-attempt-2/'), true_additional_texts_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-additional-texts/'), true_flow_dir: Path = Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-flow/'), expected_num_diagrams: int = N_DIAGRAMS_IN_BENCHMARK)

Evaluate parsed flowcharts against accepted ground-truth interpretations.

Construction loads and validates saved JSON files. Scoring methods match predicted nodes to ground-truth nodes and compare node text, labels, directed connections and additional text. Benchmarking runs locally, without opening images, changing source files or making model requests.

Parameters:
  • pred_diagrams_dir (Path | Sequence[Path]) –

    One directory containing prediction JSONs, or a sequence of directories containing separately parsed parts. Reads top-level *.json files; subdirectories are not searched. Each filename stem identifies a flowchart, and each file contains one prediction object rather than a ground-truth options list.

    Predictions must include node text. Labels, flow and additional text are optional. All JSONs within a prediction directory must contain the same parts. Separate prediction directories must contain matching filename stems and non-overlapping parts. Node-based parts must use the same node numbers for each flowchart. The benchmark joins those parts in memory; combined JSON files are not required.

  • distance_fn (DistanceFnProtocol, default: levenshtein_with_text_normalisation ) –

    Function accepting true_text and pred_text as keyword arguments and returning a finite, non-negative int or float. Either argument may be None to represent unmatched text; the supplied functions treat None as an empty string. Lower costs mean closer text matches.

    Defaults to levenshtein_with_text_normalisation(), which normalises Unicode, whitespace and selected punctuation before counting character edits. You can supply a different function to change the text comparison used for matching and scoring.

  • allow_missing_pred_diagrams (bool, default: False ) –

    Defaults to False, requiring predictions for every ground-truth flowchart. With True, evaluate only flowcharts that have predictions. Missing predictions receive no penalty and do not contribute to averages. Predictions without corresponding ground truth remain invalid.

  • true_nodes_dir (Path, default: Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-nodes-texts-numbers-row-major-column/') ) –

    Directory of ground-truth node-text JSONs. Each file contains an options list of objects with nodes; each node supplies node_number and text.

  • true_labels_dir (Path, default: Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-labels-attempt-2/') ) –

    Directory of ground-truth label JSONs. Each file contains an options list of objects with nodes; each node supplies node_number and a labels list of strings, which may be empty.

  • true_additional_texts_dir (Path, default: Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-additional-texts/') ) –

    Directory of ground-truth additional-text JSONs. Each file contains an options list of objects with an additional_texts list of strings, which may be empty.

  • true_flow_dir (Path, default: Path('./experiment_data/parsing/training-smoking-cessation/ground-truth-flow/') ) –

    Directory of ground-truth flow JSONs. Each file contains an options list of objects with nodes; each node supplies node_number and a points_to list of destination node numbers.

  • expected_num_diagrams (int, default: N_DIAGRAMS_IN_BENCHMARK ) –

    Number of flowcharts in the complete ground-truth dataset, before filtering missing predictions. Counts filename stems, not accepted interpretations. Defaults to 346, the original research dataset's size. Supply the count for your own dataset.

Attributes:
  • pred_diagrams (list[Diagram]) –

    Loaded predictions, including any parts joined from separate directories, in sorted filename-stem order.

  • true_diagrams_options_list (list[DiagramOptions]) –

    Accepted ground-truth interpretations for each evaluated flowchart, in the same filename-stem order as pred_diagrams.

  • pred_structure (PredDiagramStructure) –

    Which parts the predictions contain: node text, labels, flow and additional text.

  • capabilities (ParsingBenchmarkCapabilities) –

    Which scoring operations the predicted parts support. The available_methods tuple lists the supported method names.

Raises:
  • OSError –

    If a discovered JSON file cannot be opened.

  • TypeError –

    If a prediction or ground-truth field has an invalid type.

  • ValueError –

    If JSON is malformed, prediction directories are empty, file counts or stems disagree, Ground-Truth Options do not align, or a node or connection fails validation. Also raised for missing predicted node text, overlapping predicted parts, or inconsistent predicted parts or node numbers across their source files.

Notes

Supply all four ground-truth directories for your dataset, even when the predictions contain only some parts. The default ground-truth paths refer to the original research dataset, which is not included in a package installation. The four directories must contain matching filename stems. For each flowchart, the four options lists must have the same length; entries at the same index form one Ground-Truth Option and must agree on node numbers across node-based parts.

Scoring methods select Node Matches by minimising node-text cost. Among ties, the benchmark maximises flow similarity, then uses label cost to resolve remaining ties. Missing predicted parts are skipped. Ground-Truth Option selection follows the same order, then uses additional-text cost to resolve remaining ties; a complete tie selects the first option. All scoring methods use the available predicted parts for matching. Prediction and ground-truth node numbers do not need to match.

Text-cost methods sum matched and unmatched text costs. Flow-score methods compare directed connections after matching nodes. Results follow sorted filename-stem order, and avg_flow_jaccard() gives each evaluated flowchart equal weight. Methods requiring an unparsed part raise ValueError. A scoring method raises RuntimeError if the node matcher cannot establish an optimal solution; an invalid text-distance value raises ValueError during scoring.

Examples:

Using existing predictions and a ground-truth dataset of one flowchart:

>>> from pathlib import Path
>>> from flowde.benchmarks.parsing.parsing_bench import ParsingBenchmark
>>> truth_dir = Path("data/ground-truth")
>>> benchmark = ParsingBenchmark(
...     pred_diagrams_dir=Path("results/predictions"),
...     true_nodes_dir=truth_dir / "nodes",
...     true_labels_dir=truth_dir / "labels",
...     true_flow_dir=truth_dir / "flow",
...     true_additional_texts_dir=truth_dir / "additional_texts",
...     expected_num_diagrams=1,
... )
>>> print(benchmark.total_node_text_cost())
>>> print(benchmark.avg_flow_jaccard())

Methods

Select a method below for its parameters, return values and error conditions. Select a result type to inspect the fields and properties on each returned object. Lists follow sorted flowchart filename-stem order.

Method Required predicted parts Return value
node_matches(range_indices=None) Node text List of NodeMatches
node_matches_with_flow() Node text and flow List of NodeMatches
node_and_label_matches() Node text and labels List of NodeMatches
additional_text_matches() Node text and additional text List of TextListMatches
diagram_matches() All four parts List of DiagramMatch
total_node_text_cost() Node text float
total_label_cost() Node text and labels float
total_additional_text_cost() Node text and additional text float
all_flow_scores() Node text and flow List of FlowScores
all_flow_jaccard_scores() Node text and flow list[float]
avg_flow_jaccard() Node text and flow float

All methods use the available predicted parts to select Node Matches and a Ground-Truth Option; see how diagrams are matched. Each method calculates its results from the loaded data when called. Results are not cached between calls, and scoring makes no model requests.

node_matches

node_matches(range_indices: tuple[int | None, int | None] | None = None) -> list[NodeMatches]

Match predicted nodes to ground-truth nodes for each flowchart.

Requires predicted node text. Matching minimises node-text cost, uses flow similarity to resolve ties when flow was parsed, and then uses label cost when labels were parsed. Ground-Truth Option selection also uses additional-text cost to resolve remaining ties when that part was parsed. A complete tie selects the first Ground-Truth Option.

Parameters:
  • range_indices (tuple[int | None, int | None] | None, default: None ) –

    A (start, stop) slice of the evaluated flowcharts in sorted filename-stem order. start is included and stop is excluded. Either bound may be None; negative indices follow Python slicing rules. Defaults to None, which matches all evaluated flowcharts. An empty slice returns an empty list.

Returns:
  • list[NodeMatches] –

    One collection per selected flowchart, in sorted filename-stem order. Each collection identifies the selected Ground-Truth Option and records paired nodes, unmatched ground-truth nodes and unmatched predicted nodes, together with their node-text costs. Node numbers identify nodes within each diagram; matching does not require equal predicted and ground-truth node numbers.

Raises:
  • ValueError –

    If predicted node text is missing or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

Notes

The method uses every available predicted part for matching, even though the returned result focuses on nodes. Label matches may be populated while resolving ties; use node_and_label_matches() when you need label matches for every node.

node_matches_with_flow

node_matches_with_flow() -> list[NodeMatches]

Return Node Matches after requiring predicted flow as well as text.

Uses the same matching policy as node_matches(). The additional flow requirement ensures that each returned collection can calculate its flow_score property.

Returns:
  • list[NodeMatches] –

    One collection per evaluated flowchart, in sorted filename-stem order, including the selected Ground-Truth Option and unmatched nodes. Flow compares directed connections after applying the Node Matches.

Raises:
  • ValueError –

    If predicted node text or flow is missing, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

node_and_label_matches

node_and_label_matches() -> list[NodeMatches]

Match nodes, then match the labels belonging to each Node Match.

Requires predicted node text and labels. The benchmark selects nodes and a Ground-Truth Option using all available predicted parts, then pairs label strings within each Node Match to minimise label cost. Labels belonging to unmatched nodes count as unmatched labels.

Returns:
  • list[NodeMatches] –

    One collection per evaluated flowchart, in sorted filename-stem order. Every NodeMatch.label_matches contains a TextListMatches object, including an empty collection when both nodes have no labels. total_label_error_cost sums the collection's label costs.

Raises:
  • ValueError –

    If predicted node text or labels are missing, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

additional_text_matches

additional_text_matches() -> list[TextListMatches]

Match additional-text strings for each evaluated flowchart.

Requires predicted node text and additional text. The benchmark first selects a Ground-Truth Option using all available predicted parts, then pairs that option's additional text with the prediction to minimise additional-text cost. Strings are matched by cost, not by their positions in the input lists.

Returns:
  • list[TextListMatches] –

    One collection per flowchart, in sorted filename-stem order. Each collection contains paired and unmatched strings with their original list indices and costs. Individual records also identify the flowchart and selected Ground-Truth Option. If both diagrams have empty additional-text lists, the collection has no records and total_cost is zero.

Raises:
  • ValueError –

    If predicted node text or additional text is missing, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

diagram_matches

diagram_matches() -> list[DiagramMatch]

Match all four parsed parts for each evaluated flowchart.

Requires predicted node text, labels, flow and additional text. The benchmark selects one Ground-Truth Option per prediction and uses that option for the node, label, additional-text and flow comparisons.

Returns:
  • list[DiagramMatch] –

    One complete comparison per flowchart, in sorted filename-stem order. Each result contains the prediction, the selected ground-truth diagram, Node Matches with label matches, and additional-text matches. Text totals and flow scores are available as properties on each result.

Raises:
  • ValueError –

    If any required predicted part is missing or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

total_node_text_cost

total_node_text_cost() -> float

Sum node-text costs across all evaluated flowcharts.

Returns:
  • float –

    Total cost of paired and unmatched node text, using distance_fn. With the default distance function, the total counts character edits after text normalisation. Lower is better; the total is not an average or percentage. Requires predicted node text.

Raises:
  • ValueError –

    If predicted node text is missing or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

total_label_cost

total_label_cost() -> float

Sum label costs across all evaluated flowcharts.

Returns:
  • float –

    Total cost of paired and unmatched label strings, including labels on unmatched nodes, using distance_fn. Lower is better; the total is not an average or percentage. Requires predicted node text and labels. Node text selects the Node Matches but does not contribute to this label-only total.

Raises:
  • ValueError –

    If predicted node text or labels are missing, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

total_additional_text_cost

total_additional_text_cost() -> float

Sum additional-text costs across all evaluated flowcharts.

Returns:
  • float –

    Total cost of paired and unmatched additional-text strings, using distance_fn. Lower is better; the total is not an average or percentage. Requires predicted node text and additional text. Node text helps select the Ground-Truth Option but does not contribute to this additional-text-only total.

Raises:
  • ValueError –

    If predicted node text or additional text is missing, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

all_flow_scores

all_flow_scores() -> list[FlowScores]

Score the predicted directed connections in each flowchart.

Requires predicted node text and flow. The benchmark matches nodes before comparing connections, so predicted node numbers may differ from ground-truth node numbers. Connection direction matters: 1 -> 2 and 2 -> 1 represent different connections.

Returns:
  • list[FlowScores] –

    One score object per evaluated flowchart, in sorted filename-stem order. Each object contains correct, extra and missing connection counts, precision, recall, F1, Jaccard similarity, and the sets of extra and missing connections. Reported connections use matched ground-truth node numbers; unmatched predicted endpoints receive generated identifiers starting at 10000.

Raises:
  • ValueError –

    If predicted node text or flow is missing, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

all_flow_jaccard_scores

all_flow_jaccard_scores() -> list[float]

Return one flow Jaccard similarity score per evaluated flowchart.

Returns:
  • list[float] –

    Scores in sorted filename-stem order, calculated after matching nodes. Each score is TP / (TP + FP + FN), between 0.0 and 1.0; higher is better. A score of 1.0 means the directed connection sets agree. Two empty connection sets score 1.0. Requires predicted node text and flow.

Raises:
  • ValueError –

    If predicted node text or flow is missing, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

avg_flow_jaccard

avg_flow_jaccard() -> float

Average the flow Jaccard scores of the evaluated flowcharts.

Returns:
  • float –

    Arithmetic mean of the individual flow Jaccard scores, between 0.0 and 1.0; higher is better. Each flowchart has equal weight, regardless of its number of connections. Missing predictions excluded with allow_missing_pred_diagrams=True do not contribute to the mean. Requires predicted node text and flow.

Raises:
  • ValueError –

    If predicted node text or flow is missing, there are no evaluated flowcharts, or a text-distance value fails validation.

  • RuntimeError –

    If the node matcher cannot establish an optimal solution.

Matching result objects

These Pydantic models hold the comparisons returned by the benchmark. A match can pair a prediction with ground truth or record an unmatched node or string. The following entries describe the stored fields and computed properties.

Type Contains
NodeMatches Node Matches for one flowchart and the selected Ground-Truth Option.
NodeMatch One paired or unmatched node, with text cost and optional label matches.
TextListMatches Comparisons for one node's labels or one flowchart's additional text.
TextListMatch One paired or unmatched string, with source-list indices and cost.
DiagramMatch A complete prediction, its selected ground-truth diagram and all part comparisons.
FlowScores Directed-connection counts, metrics, and missing and extra connections.
Node An individual predicted or ground-truth node held inside a match.
Diagram The prediction or selected ground-truth diagram held inside a complete comparison.

You can inspect stored fields with model_dump() or serialise stored fields with model_dump_json(). Computed properties such as total_text_cost and flow_score are accessed directly and are not included in those dumps.

NodeMatches

Bases: BaseModel

Collect the Node Matches for one flowchart and one Ground-Truth Option.

Benchmark-generated collections include every ground-truth and predicted node exactly once, either in a pair or as an unmatched node.

Attributes:
  • matches (list[NodeMatch]) –

    Paired and unmatched node records. Inspect true_node and pred_node on each record to identify the nodes; list positions are not node numbers.

  • true_diagram_option_idx (int) –

    Index of the selected Ground-Truth Option, starting at 0.

  • parent_img_code (str) –

    Source flowchart's filename stem, taken from the first Node Match. Access raises ValueError if matches is empty.

  • true_nodes (list[Node]) –

    All present ground-truth nodes, including unmatched nodes, in the order of their records in matches.

  • pred_nodes (list[Node]) –

    All present predicted nodes, including unmatched nodes, in the order of their records in matches.

  • total_node_text_cost (float) –

    Sum of the records' node_text_cost values, including unmatched nodes.

  • total_numbers_only_node_text_cost (float) –

    Sum of the records' numbers-only text costs. Uses the existing Node Matches without choosing different pairings or a different option.

  • total_label_error_cost (float) –

    Sum of label costs across all Node Matches. Access raises ValueError if any record has label_matches=None; an empty collection costs zero.

  • flow_score (FlowScores) –

    Directed-connection scores calculated from these Node Matches. Requires a points_to list on every present node. Matched endpoints use ground-truth node numbers; unmatched predicted endpoints receive generated identifiers starting at 10000.

Raises:
  • ValidationError –

    If a field is invalid or a ground-truth or predicted node number appears in more than one record on the same side during validation.

Notes

Text costs include unmatched nodes and labels. Flow uses the same Node Matches; accessing flow_score does not select new pairings. Cost properties sum the stored records rather than rerunning model requests or matching.

get_matched_true_node
get_matched_true_node(pred_node_number: int, allow_fake_pred_node: bool = False) -> Node | None

Find the ground-truth node paired with a predicted node number.

Parameters:
  • pred_node_number (int) –

    Identifier of a predicted node recorded in matches.

  • allow_fake_pred_node (bool, default: False ) –

    Whether an absent predicted node number should return None instead of raising an error. Defaults to False. Flow scoring uses this option when looking up predicted connection endpoints.

Returns:
  • Node | None –

    The paired ground-truth node. Returns None for a recorded unmatched predicted node, or for an absent predicted node number when allow_fake_pred_node=True.

Raises:
  • ValueError –

    If pred_node_number is absent and allow_fake_pred_node=False.

get_matched_pred_node
get_matched_pred_node(true_node_number: int) -> Node | None

Find the predicted node paired with a ground-truth node number.

Parameters:
  • true_node_number (int) –

    Identifier of a ground-truth node recorded in matches.

Returns:
  • Node | None –

    The paired predicted node, or None when the recorded ground-truth node is unmatched.

Raises:
  • ValueError –

    If true_node_number does not appear in matches.

NodeMatch

Bases: BaseModel

Record one pair of corresponding nodes or one unmatched node.

A paired prediction and ground-truth node may have different text or node numbers. The benchmark stores text cost separately from any label costs.

Attributes:
  • true_node (Node | None) –

    Node from the selected Ground-Truth Option, or None when a predicted node has no ground-truth match.

  • pred_node (Node | None) –

    Predicted node, or None when a ground-truth node has no prediction.

  • node_text_cost (int | float) –

    Cost returned by the benchmark's distance_fn for the node text. An unmatched node is compared with a missing text value; built-in distance functions treat the missing value as an empty string. Excludes label and flow costs.

  • label_matches (TextListMatches | None) –

    Label comparisons for this Node Match, or None before labels have been matched. An empty collection means both nodes have no labels. Label comparisons include labels belonging to unmatched nodes.

  • match_type ({'match', 'unmatched_true', 'unmatched_pred'}) –

    "match" for two present nodes; "unmatched_true" for a ground-truth node without a prediction; "unmatched_pred" for a predicted node without ground truth. A pair need not have zero text cost.

  • parent_img_code (str) –

    Source flowchart's filename stem, taken from the ground-truth node when present, otherwise from the predicted node.

  • numbers_only_node_text_cost (float) –

    Cost from comparing only the numbers in the two nodes' text with number_only_levenshtein(). Uses this existing Node Match without changing the nodes or their pairing.

Raises:
  • ValidationError –

    If a field is invalid, both nodes are absent, or two present nodes have different parent_img_code values.

TextListMatches

Bases: BaseModel

Collect comparisons for one node's labels or a diagram's additional text.

The benchmark matches strings to minimise total cost rather than pairing strings by their list positions. Each record retains the original indices so the source strings can be identified.

Attributes:
  • matches (list[TextListMatch]) –

    Paired and unmatched string records. Benchmark-generated collections account for every string in both source lists. Two empty source lists produce an empty collection. Use each record's true_index and pred_index to identify its original positions.

  • total_cost (float) –

    Sum of every record's cost, including unmatched strings. An empty collection has cost zero. The property reports a total, not an average.

Notes

This class stores the supplied records and computes their total; the class does not perform string matching when constructed. The container does not store source lists or flowchart metadata. For additional text, the individual records hold the flowchart and Ground-Truth Option IDs.

TextListMatch

Bases: BaseModel

Record one paired or unmatched label or additional-text string.

A pair records which predicted string represents which ground-truth string, even when their text differs. An unmatched record contains text on only one side.

Attributes:
  • true_index (int | None) –

    Zero-based position in the ground-truth text list. None when a predicted string has no ground-truth match.

  • pred_index (int | None) –

    Zero-based position in the predicted text list. None when a ground-truth string has no predicted match.

  • true_text (str | None) –

    Original ground-truth string, before normalisation; None exactly when true_index is None. Present strings must be non-empty.

  • pred_text (str | None) –

    Original predicted string, before normalisation; None exactly when pred_index is None. Present strings must be non-empty.

  • parent_img_code (str | None) –

    Source flowchart's filename stem. Additional-text matching sets this field; label matching leaves the default None because the containing NodeMatch identifies the flowchart.

  • true_option_idx (int | None) –

    Selected Ground-Truth Option index, starting at 0. Additional-text matching sets this field; label matching leaves the default None.

  • cost (float) –

    Distance-function cost for this string pair or unmatched string. An unmatched string is compared with None; the built-in distance functions treat the missing side as an empty string.

  • match_type ({'match', 'unmatched_true', 'unmatched_pred'}) –

    "match" when both strings are present; "unmatched_true" for a ground-truth string without a prediction; "unmatched_pred" for a predicted string without ground truth. "match" does not imply identical text or zero cost.

Raises:
  • ValidationError –

    If a field is invalid, an index and its text disagree about whether the corresponding side is absent, or both sides are absent.

DiagramMatch

Bases: BaseModel

Collect a complete prediction's comparison with one Ground-Truth Option.

Both diagrams contain node text, labels, flow and additional text. The result exposes the loaded diagrams, their matching records, and the text and flow scores calculated from those records.

Attributes:
  • node_matches (NodeMatches) –

    Every paired and unmatched node, with label matches for each record. true_diagram_option_idx identifies the selected Ground-Truth Option.

  • additional_text_matches (TextListMatches) –

    Paired and unmatched additional-text strings from both diagrams.

  • pred_diagram (Diagram) –

    Complete predicted diagram, retaining its original node numbers and text.

  • true_diagram (Diagram) –

    Complete ground-truth diagram selected from the accepted options. true_option_idx is its zero-based option index.

  • total_node_text_cost (float) –

    Sum of paired and unmatched node-text costs.

  • total_label_error_cost (float) –

    Sum of paired and unmatched label costs, including labels on unmatched nodes.

  • total_additional_text_cost (float) –

    Sum of paired and unmatched additional-text costs.

  • total_text_cost (float) –

    Sum of node-text, label and additional-text costs. Excludes flow scores.

  • flow_score (FlowScores) –

    Directed-connection scores after applying node_matches.

Raises:
  • ValidationError –

    If a field is invalid, either diagram is incomplete, the diagrams refer to different flowcharts or have incorrect diagram_type values, or the node and additional-text records do not account for the contents of the compared diagrams.

Flow scores

FlowScores

Bases: BaseModel

Store directed-connection counts and scores for one matched flowchart.

The benchmark calculates these fields after matching predicted nodes to ground-truth nodes. Connection direction matters: 1 -> 2 differs from 2 -> 1.

Attributes:
  • tp (int) –

    Number of connections present in both the prediction and ground truth.

  • fp (int) –

    Number of predicted connections absent from ground truth.

  • fn (int) –

    Number of ground-truth connections absent from the prediction.

  • precision (float) –

    Correct connections divided by predicted connections: tp / (tp + fp). Zero when the prediction has no connections.

  • recall (float) –

    Correct connections divided by ground-truth connections: tp / (tp + fn). Zero when ground truth has no connections.

  • f1 (float) –

    Harmonic mean of precision and recall. Zero when both are zero.

  • jaccard (float) –

    Correct connections divided by all distinct connections in either diagram: tp / (tp + fp + fn). Ranges from 0.0 to 1.0; higher is better. Two empty connection sets score 1.0.

  • missing_edges (set[tuple[int, int]]) –

    Missing ground-truth connections as (source, destination) pairs of ground-truth node numbers.

  • extra_edges (set[tuple[int, int]]) –

    Predicted connections absent from ground truth, expressed using matched ground-truth node numbers. Unmatched predicted endpoints use generated identifiers starting at 10000, not original node numbers.

Notes

A FlowScores object stores supplied values; constructing the object does not calculate metrics from tp, fp and fn. Benchmark methods calculate the metrics before constructing the object.

Nodes and diagrams

These benchmark models preserve the source text and node numbers alongside the metadata identifying the flowchart and Ground-Truth Option.

Node

Bases: BaseModel

Represent one predicted or ground-truth node inside a parsing benchmark.

Benchmark loaders attach the flowchart identifier and Ground-Truth Option index to the node fields read from JSON. Node numbers identify nodes within a diagram; predicted and ground-truth versions of a node may have different numbers.

Attributes:
  • node_number (int) –

    Node identifier. Must be less than 1000; the containing Diagram checks uniqueness and applies the ground-truth numbering rules.

  • text (str) –

    Non-empty node text as loaded, before any distance-function normalisation.

  • labels (list[str] | None) –

    Non-empty label strings. [] means labels were parsed and none were found; None, the default, means labels were not parsed.

  • points_to (list[int] | None) –

    Destination node numbers for outgoing connections. [] means flow was parsed and the node has no outgoing connections; None, the default, means flow was not parsed.

  • diagram_type ({'pred', 'true'}) –

    Whether the node belongs to a prediction or to ground truth.

  • true_option_idx (int | None) –

    Zero-based Ground-Truth Option index. Required for a ground-truth node; must be None for a predicted node. Defaults to None.

  • parent_img_code (str) –

    Filename stem identifying the source flowchart.

  • text_done (bool) –

    Whether node text is present; always True for a valid Node.

  • labels_done (bool) –

    Whether labels is present, including an empty list.

  • flow_done (bool) –

    Whether points_to is present, including an empty list.

  • is_complete (bool) –

    Whether text, labels and flow are all present.

Raises:
  • ValidationError –

    If a field has an invalid type, text or a label string is empty, node_number is at least 1000, or true_option_idx conflicts with diagram_type.

Diagram

Bases: BaseModel

Represent one prediction or one accepted ground-truth interpretation.

A DiagramMatch exposes the compared diagrams through pred_diagram and true_diagram. Diagram fields retain the original text and node numbers; matching and text normalisation do not rewrite the loaded diagrams.

Attributes:
  • nodes (list[Node] | None) –

    Non-empty list of nodes, or None when no node-based parts were supplied. Defaults to None. Parsing benchmarks require predicted node text, so diagrams returned by benchmark methods contain nodes.

  • additional_texts (list[str] | None) –

    Non-empty strings outside the nodes. [] means additional text was parsed and none was found; None, the default, means the part was not parsed.

  • parent_img_code (str) –

    Filename stem identifying the source flowchart.

  • diagram_type ({'pred', 'true'}) –

    Whether the diagram is a prediction or a Ground-Truth Option.

  • true_option_idx (int | None) –

    Zero-based Ground-Truth Option index for a ground-truth diagram; None for a predicted diagram.

  • text_done (bool) –

    Whether nodes and their text are present.

  • labels_done (bool) –

    Whether every node has a labels list, including empty lists.

  • flow_done (bool) –

    Whether every node has a points_to list, including empty lists.

  • additional_texts_done (bool) –

    Whether additional_texts is present, including an empty list.

  • is_complete (bool) –

    Whether node text, labels, flow and additional text are all present.

Raises:
  • ValidationError –

    If fields are invalid, node numbers repeat, connections refer to absent nodes, or node metadata disagrees with the diagram. Also raised if only some nodes supply labels or flow, or a ground-truth diagram is incomplete, has non-sequential node numbers, or contains a node absent from every connection.

Notes

Ground-truth nodes must appear in consecutive node-number order starting at 1. Predictions may use different numbering. Ground-truth diagrams must contain all four parts; predictions may omit labels, flow or additional text.

Text distances

Import the following functions from flowde.benchmarks.parsing.text_distance_fns.levenshtein_fn.

levenshtein_with_text_normalisation

levenshtein_with_text_normalisation(true_text: str | None, pred_text: str | None) -> int

Count character edits after normalising both input strings.

This is the default text distance used by ParsingBenchmark. Normalisation reduces differences caused by formatting and alternative character forms before Levenshtein distance counts insertions, deletions and substitutions.

Parameters:
  • true_text (str | None) –

    Ground-truth text. None represents unmatched predicted text and is treated as an empty string.

  • pred_text (str | None) –

    Predicted text. None represents missing predicted text and is treated as an empty string.

Returns:
  • int –

    Minimum number of single-character edits between the normalised strings. 0 means the normalised strings are identical. The result is an edit count, not a score scaled between zero and one.

Raises:
  • ValueError –

    If both true_text and pred_text are None.

Notes

Both strings receive the same normalisation:

  • Apply Unicode NFC so equivalent composed and decomposed characters match.
  • Standardise line endings, convert tabs to spaces, collapse repeated ordinary spaces and newlines, and strip surrounding whitespace.
  • Standardise supported bullet, dash, quotation-mark and caret characters. A line beginning with o or O followed by whitespace becomes a bullet.
  • Standardise spacing around punctuation, brackets, equals signs, plus signs and hyphens between word characters, including digits.
  • Collapse repeated HTML line-break tags to one <br> tag.
  • Convert superscript ordinal suffixes following digits, such as 1ˢᵗ to 1st.

Normalisation does not lowercase the text or remove accents. Single newlines remain distinct from spaces unless a punctuation-spacing rule removes the newline. HTML break tags are not converted to newlines.

Examples:

>>> levenshtein_with_text_normalisation("n=10", "n = 10")
0
>>> levenshtein_with_text_normalisation("n = 10", "n = 11")
1
>>> levenshtein_with_text_normalisation(None, " ABC ")
3

ParsingBenchmark uses this function when you omit distance_fn. You can still pass levenshtein_fn() for raw character comparisons, or another distance function.

levenshtein_fn

levenshtein_fn(true_text: str | None, pred_text: str | None) -> int

Counts the minimum number of single-character insertions, deletions and substitutions. Treats a missing string as empty; two missing strings are invalid.

number_only_levenshtein

number_only_levenshtein(true_text: str | None, pred_text: str | None) -> int

Compares the extracted numbers in the two strings. Use this for a separate number-focused check, not as a measure of whether the surrounding words agree. NodeMatches.total_numbers_only_node_text_cost provides this check for the already selected node matches.

DistanceFnProtocol

Bases: Protocol

Describe a callable that calculates a cost between true and predicted text.

Pass a function with this interface as distance_fn to ParsingBenchmark to compare node text, labels and additional text. Lower costs must indicate closer matches because the benchmark minimises text costs when matching predictions to ground truth.

Notes

This protocol describes the callable interface; it does not calculate costs or validate returned values. A function or callable object can satisfy the interface without inheriting from DistanceFnProtocol.

__call__

__call__(*, true_text: str | None, pred_text: str | None) -> int | float

Calculate the comparison cost for one pair of text values.

Parameters:
  • true_text (str | None) –

    Ground-truth text. The benchmark supplies this argument by keyword. None represents predicted text with no ground-truth match.

  • pred_text (str | None) –

    Predicted text. The benchmark supplies this argument by keyword. None represents ground-truth text with no prediction.

Returns:
  • int | float –

    Finite, non-negative comparison cost. Lower values must indicate closer text matches. Return a cost rather than a similarity score where higher values indicate closer matches.

Notes

The function must handle either argument being None and assign a cost to the unmatched text. Flowde's built-in distance functions treat None as an empty string. A custom function can choose another penalty for unmatched text. The benchmark does not compare two None values.