Pipeline functions

All four stages save run state in the .flowde directory inside save_dir. Each stage needs its own output directory. Do not put source files or unrelated results in that directory. See stopping and resuming for the complete recovery rules.

Shared parameters

Parameter Meaning
save_dir Directory for this run's results and .flowde records. Created when needed.
on_existing "error" (default) rejects an existing run; "resume" restores a compatible run; "overwrite" starts a fresh run and replaces the previous run's tracked outputs. Overwrite also works with a new or empty directory.
n_jobs Extraction only: number of worker processes, default 1.
max_concurrent_jobs Classification, rotation and parsing: maximum image jobs in a separate process's thread pool, default 10. Must be a positive integer.
range_indices On classification, rotation and parsing, a (start, stop) slice of sorted input paths, with stop excluded. None selects all paths. Extraction does not take this parameter.
show_usage On classification, rotation and parsing, show the token and estimated cost summary. False hides the summary without disabling saved usage reports.

max_concurrent_jobs can change when resuming a run.

Image functions return results for the selected inputs in sorted input order. Resuming an overlapping slice restores completed answers within that slice. Answers outside the current slice remain saved, but are not included in the returned list.

Input filename stems must be unique ignoring case across the run, including inputs from earlier resumed calls. For example, diagram.png and DIAGRAM.jpg cannot belong to the same run, even in separate calls.

Saved run state

Keep the whole .flowde directory with the outputs. Its files have these roles:

Path inside .flowde Contents
run_metadata.state Format version, pipeline stage, settings and input paths.
input_records/<stem>.state One input's result, reported usage, error, completion status and output fingerprints.
shared_output_fingerprints.state Fingerprints for classifications.json or rotations.json; empty for parsing and extraction.
run.lock Lock that prevents concurrent runs from writing to the same output directory.

Progress updates rewrite only the affected input record. Run metadata is written when inputs are first registered or added on resume. Parsing saves each result before publishing its output JSON, so resume can recover a saved answer without another parsing request. Classification and rotation still rewrite their combined output JSON and its fingerprints as results are published.

On resume and overwrite, Flowde validates all expected internal records, including records outside the selected slice. Missing or invalid records cause an error; Flowde does not reconstruct them from output files. This differs from a missing public output: saved image results can recreate missing JSONs or image copies, while missing extracted PNGs require re-extracting the source PDF.

Extraction

extract_imgs

extract_imgs(pdf_dir: Path, save_dir: Path, extract_fn: ExtractImgsFunction, n_jobs: int = 1, *, on_existing: ExistingRun = 'error') -> None

Extract sorted top-level PDFs, saving complete image sets and resumable state.

Parameters:
  • pdf_dir (Path) –

    Directory containing input PDFs. Subdirectories are not searched. Filename stems must be unique ignoring case across the run, including PDFs from earlier resumed calls.

  • save_dir (Path) –

    Dedicated output directory containing PNGs and .flowde run metadata.

  • extract_fn (ExtractImgsFunction) –

    Function accepting pdf_path and save_dir as keyword arguments, both containing Path objects. Flowde supplies a temporary directory for one PDF as the extractor's save_dir.

    The extractor must save valid PNG files with the .png extension directly inside the supplied directory, without subdirectories, symbolic links or other files. Output filenames must be unique across all PDFs in the run, including when compared without letter case. The extractor must finish writing and close all output files before returning None. Flowde checks the PNGs and copies them into the run's output directory only after the extractor finishes successfully.

    Built-in factories declare their settings automatically. Custom extractors must declare their settings with model_function().

  • n_jobs (int, default: 1 ) –

    Number of PDFs processed concurrently. Defaults to 1.

  • on_existing ((error, resume, overwrite), default: "error" ) –

    How to handle an existing run in save_dir. Defaults to "error".

    • "error": start in a new or empty directory; reject existing work.
    • "resume": require a saved extraction run with matching settings. Skip completed PDFs with intact PNGs and re-extract unfinished PDFs or those with missing PNGs. Edited saved PNGs cause an error.
    • "overwrite": remove the previous run's tracked outputs and saved state, then start a new run. Also works in a new or empty directory. Unrelated files in an existing output directory cause an error.

    Resume and overwrite require all internal records. Missing or invalid records raise an error.

Returns:
  • None –

    Extracted images are written into save_dir.

Raises:
  • FileExistsError –

    If save_dir contains existing work and on_existing="error".

  • ValueError –

    If pdf_dir is invalid, no PDFs are found, filename stems conflict ignoring case, inputs are inside save_dir, extraction outputs are invalid or conflict, or the saved run fails compatibility or integrity checks.

  • RuntimeError –

    If another Flowde call is already using the same save_dir.

Notes

Flowde saves settings in .flowde/run_metadata.state and each PDF's progress and PNG fingerprints in .flowde/input_records/<stem>.state, including PDFs that produce no images. Progress updates rewrite only that PDF's record. Keep the whole .flowde directory with the outputs for resume and overwrite. The records do not contain image data; missing PNGs require re-extraction.

Parameter Meaning
pdf_dir Directory containing top-level *.pdf files with distinct filename stems, including when compared without letter case.
extract_fn Callable accepting pdf_path and save_dir keyword arguments. The callable writes top-level PNGs into the supplied temporary directory.

Returns None. After an entire PDF succeeds, Flowde publishes that PDF's PNGs into the run's save_dir. The run record includes successful PDFs that produced no images. Missing PNGs cause the whole source PDF to be extracted again on resume; modified saved PNGs cause an error.

For the Paddle factory's crop names and settings, see image extraction.

Classification

classify_imgs

classify_imgs(classify_fn: ClassificationFunction[LabelType], img_dir: Path, save_dir: Path, *, positive_classes: set[ClassificationLabel] | None = None, range_indices: tuple[int | None, int | None] | None = None, max_concurrent_jobs: int = 10, show_usage: bool = True, on_existing: ExistingRun = 'error') -> list[LabelType]

Classify PNG images in a directory and save their labels in a resumable run.

Parameters:
  • classify_fn (ClassificationFunction[LabelType]) –

    Function accepting an image Path as its first positional argument and returning a str, int or bool label. The function returns the label itself, not a dictionary or Pydantic model. Built-in factories declare their settings; custom functions must declare their settings with model_function(). If the classifier exposes a result_structure, returned labels are validated against that schema.

  • img_dir (Path) –

    Directory containing the input *.png files. Subdirectories are not searched. Image paths are sorted before applying range_indices. Selected filename stems must be unique ignoring case across the run, including images from earlier resumed calls.

  • save_dir (Path) –

    Dedicated output directory, created if needed. Labels are saved in classifications.json, run metadata in .flowde, and optional image copies in positive_images. Input images and files referenced by the classifier's declared settings must be outside this directory.

  • positive_classes (set[ClassificationLabel] | None, default: None ) –

    Labels whose images are copied into save_dir / "positive_images", preserving filenames and leaving the original images unchanged. For example, {1} copies images labelled 1. Defaults to None, which saves labels without copying images. Must remain unchanged when resuming the run.

  • range_indices (tuple[int | None, int | None] | None, default: None ) –

    A (start, stop) slice of the sorted image paths. start is included and stop is excluded; (0, 10) selects up to the first ten images. Either bound can be None, and negative indices follow Python slicing rules. Defaults to None, which selects all matching images. An empty selection raises an error.

  • max_concurrent_jobs (int, default: 10 ) –

    Maximum number of images handled concurrently by threads in a separate worker process, by default 10. Must be a positive integer; 1 handles one image at a time. Custom functions must support concurrent calls when this exceeds 1. Even at 1, worker mutations do not update objects in the caller; see custom functions. This setting can change when resuming a run.

  • show_usage (bool, default: True ) –

    Whether to display token usage and estimated costs reported by the classifier. Defaults to True. False hides those figures while retaining the progress display and saved usage reports. This value can change when resuming a run.

  • on_existing ((error, resume, overwrite), default: "error" ) –

    How to handle an existing run in save_dir. Defaults to "error".

    • "error": start in a new or empty directory; reject existing work.
    • "resume": require a saved classification run with matching classifier settings and positive_classes. Reuse saved labels for unchanged inputs and classify selected images without saved labels.
    • "overwrite": remove the previous run's tracked outputs and saved state, then start a new run. Also works in a new or empty directory. Unrelated files in an existing output directory cause an error.

    Resume and overwrite require all internal records. Missing or invalid records raise an error.

Returns:
  • list[LabelType] –

    One label per image selected by this call, in sorted image-path order. Includes labels restored from a previous run. With range_indices, the returned list covers only the selected slice; labels for earlier slices remain saved in classifications.json.

Raises:
  • FileExistsError –

    If save_dir contains existing work and on_existing="error".

  • ValueError –

    If img_dir is missing or is not a directory, no PNGs are selected, filename stems conflict ignoring case, inputs are inside save_dir, or the saved run fails compatibility or integrity checks.

  • TypeError –

    If the classifier returns a label that is not a string, integer or boolean, including None.

  • RuntimeError –

    If another Flowde call is already using the same save_dir.

Notes

classifications.json contains objects with img_path and label fields. The JSON file accumulates labels across resumed calls and orders entries by resolved input paths. Flowde saves settings in .flowde/run_metadata.state and each image's label, usage and progress in .flowde/input_records. Keep the whole .flowde directory with the outputs for resume and overwrite.

On resume, missing classifications.json or selected positive_images copies are recreated from saved labels and unchanged input images without calling the classifier again. Manually edited saved outputs raise an error.

See the classification guide for worked examples.

Rotation

rotate_imgs

rotate_imgs(classify_fn: ClassificationFunction[RotationLabel], img_dir: Path, save_dir: Path, *, range_indices: tuple[int | None, int | None] | None = None, max_concurrent_jobs: int = 10, show_usage: bool = True, on_existing: ExistingRun = 'error') -> list[RotationLabel]

Rotate PNG images in a directory and save corrected copies in a resumable run.

Parameters:
  • classify_fn (ClassificationFunction[RotationLabel]) –

    Function accepting an image Path as its first positional argument and returning a clockwise correction in degrees as an int: 0, 90, 180 or 270. The function returns the angle itself, not a dictionary or Pydantic model, and must raise an exception if angle prediction fails. Returning None raises an error.

    The classifier must leave the input image unchanged. Flowde applies the returned angle and saves the corrected copy. Built-in factories declare their settings; custom functions must declare their settings with model_function().

  • img_dir (Path) –

    Directory containing the original *.png files. Subdirectories are not searched. Image paths are sorted before applying range_indices. Selected filename stems must be unique ignoring case across the run, including images from earlier resumed calls.

  • save_dir (Path) –

    Dedicated output directory, created if needed. Correction angles are saved in rotations.json, corrected copies in rotated_images, and run metadata in .flowde. Original images remain unchanged. Input images and files referenced by the classifier's declared settings must be outside this directory.

  • range_indices (tuple[int | None, int | None] | None, default: None ) –

    A (start, stop) slice of the sorted image paths. start is included and stop is excluded; (0, 10) selects up to the first ten images. Either bound can be None, and negative indices follow Python slicing rules. Defaults to None, which selects all matching images. An empty selection raises an error.

  • max_concurrent_jobs (int, default: 10 ) –

    Maximum number of images handled concurrently by threads in a separate worker process, by default 10. Must be a positive integer; 1 handles one image at a time. Custom functions must support concurrent calls when this exceeds 1. Even at 1, worker mutations do not update objects in the caller; see custom functions. This setting can change when resuming a run.

  • show_usage (bool, default: True ) –

    Whether to display token usage and estimated costs reported by the classifier. Defaults to True. False hides those figures while retaining the progress display and saved usage reports. This value can change when resuming a run.

  • on_existing ((error, resume, overwrite), default: "error" ) –

    How to handle an existing run in save_dir. Defaults to "error".

    • "error": start in a new or empty directory; reject existing work.
    • "resume": require a saved rotation run with matching classifier settings. Reuse saved angles for unchanged inputs and predict angles for selected images without saved angles.
    • "overwrite": remove the previous run's tracked outputs and saved state, then start a new run. Also works in a new or empty directory. Unrelated files in an existing output directory cause an error.

    Resume and overwrite require all internal records. Missing or invalid records raise an error.

Returns:
  • list[RotationLabel] –

    One clockwise correction angle per image selected by this call, in sorted image-path order. Includes angles restored from a previous run. With range_indices, the returned list covers only the selected slice; angles for earlier slices remain saved in rotations.json.

Raises:
  • FileExistsError –

    If save_dir contains existing work and on_existing="error".

  • ValueError –

    If img_dir is missing or is not a directory, no PNGs are selected, filename stems conflict ignoring case, an angle is unsupported, inputs are inside save_dir, or the saved run fails compatibility or integrity checks.

  • TypeError –

    If the classifier returns None instead of an angle.

  • RuntimeError –

    If another Flowde call is already using the same save_dir.

Notes

rotations.json contains objects with img_path and label fields, where label is the clockwise correction angle. The JSON file accumulates angles across resumed calls and orders entries by resolved input paths.

Flowde saves settings in .flowde/run_metadata.state and each image's angle, usage and progress in .flowde/input_records. Keep the whole .flowde directory with the outputs for resume and overwrite.

Every successfully processed image has a copy in rotated_images, including images with a zero-degree correction. Copies retain their filenames and expand to fit the rotated image without cropping.

On resume, missing rotations.json or selected rotated_images copies are recreated from saved angles and unchanged original images without another prediction. Corrected copies are always made from the original inputs, so resuming does not rotate an already-corrected copy again. Edited previously processed input images or saved outputs cause an error.

See the rotation guide for worked examples.

Parsing

parse_imgs

parse_imgs(parse_fn: ParsingFunction, img_dir: Path, save_dir: Path, nodes_dir: Path | None = None, labels_dir: Path | None = None, additional_texts_dir: Path | None = None, flow_dir: Path | None = None, range_indices: tuple[int | None, int | None] | None = None, max_concurrent_jobs: int = 10, img_extensions: set[str] | None = None, *, show_usage: bool = True, on_existing: ExistingRun = 'error') -> list[BaseModel]

Parse images in a directory into JSON files, optionally using saved context.

Parameters:
  • parse_fn (ParsingFunction) –

    Function accepting an img_path keyword argument and, when context is supplied, a partial_flowchart keyword argument containing a Pydantic model. img_path is a Path to the image being parsed; the partial model contains previously parsed parts for that same image. Without saved context, Flowde passes only img_path.

    The function must return an instance of the Pydantic class exposed as parse_fn.result_structure, not a dictionary, JSON string or None. The class can describe standard flowchart parts or a custom output format. The parser must raise an exception if parsing fails. Flowde saves each returned model as a same-stem JSON file in save_dir.

    Built-in factories declare their settings; custom functions must declare their settings and result class with model_function().

  • img_dir (Path) –

    Directory containing input images. Subdirectories are not searched. Matching paths are sorted before applying range_indices. Filename stems must be unique across all matching images, including different extensions, because each stem determines an output JSON filename. Selected stems must also remain unique ignoring case across the run, including images from earlier resumed calls.

  • save_dir (Path) –

    Dedicated output directory, created if needed. Each image produces a JSON file with the same stem, such as diagram.png producing diagram.json. Run metadata is saved in .flowde. Input images, partial JSONs and files referenced by the parser's declared settings must be outside this directory.

  • nodes_dir (Path | None, default: None ) –

    Directory of node-text JSONs supplied as context alongside the images. Each JSON contains nodes with node_number and text fields. Defaults to None, meaning no saved context is supplied. Required if labels_dir, additional_texts_dir or flow_dir is provided.

  • labels_dir (Path | None, default: None ) –

    Directory of label JSONs to add to the context from nodes_dir. Each JSON contains nodes with node_number and labels fields. Defaults to None, meaning no saved labels are included in the context.

  • additional_texts_dir (Path | None, default: None ) –

    Directory of additional-text JSONs to add to the context from nodes_dir. Each JSON contains an additional_texts list. Defaults to None, meaning no saved additional text is included in the context.

  • flow_dir (Path | None, default: None ) –

    Directory of connection JSONs to add to the context from nodes_dir. Each JSON contains nodes with node_number and points_to fields. Defaults to None, meaning no saved connections are included in the context.

  • range_indices (tuple[int | None, int | None] | None, default: None ) –

    A (start, stop) slice of the sorted image paths. start is included and stop is excluded; (0, 10) selects up to the first ten images. Either bound can be None, and negative indices follow Python slicing rules. Defaults to None, which selects all matching images. The same slice is applied to supplied context files after checking their filenames against all matching images. An empty selection raises an error.

  • max_concurrent_jobs (int, default: 10 ) –

    Maximum number of images handled concurrently by threads in a separate worker process, by default 10. Must be a positive integer; 1 handles one image at a time. Custom functions must support concurrent calls when this exceeds 1. Even at 1, worker mutations do not update objects in the caller; see custom functions. This setting can change when resuming a run.

  • img_extensions (set[str] | None, default: None ) –

    Image extensions to select, without leading dots, such as {"png", "jpg", "webp"}. Defaults to None, which selects PNG files. The supplied parser must support the selected image formats.

  • show_usage (bool, default: True ) –

    Whether to display token usage and estimated costs reported by the parser. Defaults to True. False hides those figures while retaining the progress display and saved usage reports. This value can change when resuming a run.

  • on_existing ((error, resume, overwrite), default: "error" ) –

    How to handle an existing run in save_dir. Defaults to "error".

    • "error": start in a new or empty directory; reject existing work.
    • "resume": require a saved parsing run with matching parser settings. Restore saved results for unchanged images and assembled partial context, and parse selected images without saved results.
    • "overwrite": remove the previous run's tracked outputs and saved state, then start a new run. Also works in a new or empty directory. Unrelated files in an existing output directory cause an error.

    Resume and overwrite require all internal records. Missing or invalid records raise an error.

Returns:
  • list[BaseModel] –

    One Pydantic result per image selected by this call, in sorted image-path order. Each result uses parse_fn.result_structure. Includes results restored from a previous run. With range_indices, the returned list covers only the selected slice; JSON files for earlier slices remain saved in save_dir.

Raises:
  • FileExistsError –

    If save_dir contains existing work and on_existing="error".

  • ValueError –

    If img_dir is invalid, no images are selected, image stems conflict (including case-only differences across the run), context files do not match the images or required part schemas, node numbers cannot be joined, inputs are inside save_dir, or the saved run fails compatibility or integrity checks.

  • TypeError –

    If parse_fn.result_structure is not a Pydantic class or the parser returns a value that is not an instance of that class, including None.

  • RuntimeError –

    If another Flowde call is already using the same save_dir.

Notes

Each supplied context directory must contain exactly one top-level *.json file per matching input image, with the same stem and no extra JSON files. Filename and file-count checks cover all matching images before slicing, including images outside range_indices.

Context files contain the fields for their individual parts, without the benchmark ground truth's options wrapper. Selected node-text files must contain at least one node, numbered consecutively from 1. Selected label and flow files must contain the same node numbers as the node-text file. The parts are joined into the partial_flowchart passed to the parser.

Context files supply input to the parser; they are not automatically merged into its output. The output fields are defined by parse_fn.result_structure. For built-in parsers, the factory's parts_to_parse or result_structure argument chooses those fields.

Flowde saves settings in .flowde/run_metadata.state and each image's result, usage and progress in .flowde/input_records/<stem>.state. Progress updates rewrite only that image's record, without rewriting other images' results or the run metadata. Keep the whole .flowde directory with the outputs.

Flowde records each result before writing its output JSON. On resume, missing output JSONs are recreated from these records without another parsing request. Edited saved outputs or changes to previously parsed input images or their assembled partial context cause an error. Missing or invalid internal records cause an error even when the output JSONs still exist.

See the parsing guide for full, partial and custom-schema examples.