Pipeline functions
All four stages save run state in the .flowde directory inside
save_dir. Each stage needs its own output directory. Do not put source files
or unrelated results in that directory. See
stopping and resuming for the complete recovery
rules.
Shared parameters
| Parameter |
Meaning |
save_dir |
Directory for this run's results and .flowde records. Created when needed. |
on_existing |
"error" (default) rejects an existing run; "resume" restores a compatible run; "overwrite" starts a fresh run and replaces the previous run's tracked outputs. Overwrite also works with a new or empty directory. |
n_jobs |
Extraction only: number of worker processes, default 1. |
max_concurrent_jobs |
Classification, rotation and parsing: maximum image jobs in a separate process's thread pool, default 10. Must be a positive integer. |
range_indices |
On classification, rotation and parsing, a (start, stop) slice of sorted input paths, with stop excluded. None selects all paths. Extraction does not take this parameter. |
show_usage |
On classification, rotation and parsing, show the token and estimated cost summary. False hides the summary without disabling saved usage reports. |
max_concurrent_jobs can change when resuming a run.
Image functions return results for the selected inputs in sorted input order.
Resuming an overlapping slice restores completed answers within that slice.
Answers outside the current slice remain saved, but are not included in the
returned list.
Input filename stems must be unique ignoring case across the run, including
inputs from earlier resumed calls. For example, diagram.png and DIAGRAM.jpg
cannot belong to the same run, even in separate calls.
Saved run state
Keep the whole .flowde directory with the outputs. Its files have these roles:
Path inside .flowde |
Contents |
run_metadata.state |
Format version, pipeline stage, settings and input paths. |
input_records/<stem>.state |
One input's result, reported usage, error, completion status and output fingerprints. |
shared_output_fingerprints.state |
Fingerprints for classifications.json or rotations.json; empty for parsing and extraction. |
run.lock |
Lock that prevents concurrent runs from writing to the same output directory. |
Progress updates rewrite only the affected input record. Run metadata is written
when inputs are first registered or added on resume. Parsing saves each result
before publishing its output JSON, so resume can recover a saved answer without
another parsing request. Classification and rotation still rewrite their combined
output JSON and its fingerprints as results are published.
On resume and overwrite, Flowde validates all expected internal records,
including records outside the selected slice. Missing or invalid records cause
an error; Flowde does not reconstruct them from output files. This differs from
a missing public output: saved image results can recreate missing JSONs or image
copies, while missing extracted PNGs require re-extracting the source PDF.
extract_imgs(pdf_dir: Path, save_dir: Path, extract_fn: ExtractImgsFunction, n_jobs: int = 1, *, on_existing: ExistingRun = 'error') -> None
Extract sorted top-level PDFs, saving complete image sets and resumable state.
| Parameters: |
-
pdf_dir
(Path)
–
Directory containing input PDFs. Subdirectories are not searched.
Filename stems must be unique ignoring case across the run, including
PDFs from earlier resumed calls.
-
save_dir
(Path)
–
Dedicated output directory containing PNGs and .flowde run metadata.
-
extract_fn
(ExtractImgsFunction)
–
Function accepting pdf_path and save_dir as keyword arguments,
both containing Path objects. Flowde supplies a temporary directory
for one PDF as the extractor's save_dir.
The extractor must save valid PNG files with the .png extension
directly inside the supplied directory, without subdirectories,
symbolic links or other files. Output filenames must be unique across
all PDFs in the run, including when compared without letter case.
The extractor must finish writing and close all output files before
returning None. Flowde checks the PNGs and copies them into the run's
output directory only after the extractor finishes successfully.
Built-in factories declare their settings automatically. Custom
extractors must declare their settings with
model_function().
-
n_jobs
(int, default:
1
)
–
Number of PDFs processed concurrently. Defaults to 1.
-
on_existing
((error, resume, overwrite), default:
"error"
)
–
How to handle an existing run in save_dir. Defaults to "error".
"error": start in a new or empty directory; reject existing work.
"resume": require a saved extraction run with matching settings.
Skip completed PDFs with intact PNGs and re-extract unfinished PDFs
or those with missing PNGs. Edited saved PNGs cause an error.
"overwrite": remove the previous run's tracked outputs and saved state,
then start a new run. Also works in a new or empty directory.
Unrelated files in an existing output directory cause an error.
Resume and overwrite require all internal records. Missing or invalid
records raise an error.
|
| Returns: |
-
None
–
Extracted images are written into save_dir.
|
| Raises: |
-
FileExistsError
–
If save_dir contains existing work and on_existing="error".
-
ValueError
–
If pdf_dir is invalid, no PDFs are found, filename stems conflict
ignoring case, inputs are inside save_dir, extraction outputs are
invalid or conflict, or the saved run fails compatibility or integrity
checks.
-
RuntimeError
–
If another Flowde call is already using the same save_dir.
|
Notes
Flowde saves settings in .flowde/run_metadata.state and each PDF's progress
and PNG fingerprints in .flowde/input_records/<stem>.state, including PDFs
that produce no images. Progress updates rewrite only that PDF's record.
Keep the whole .flowde directory with the outputs for resume and overwrite.
The records do not contain image data; missing PNGs require re-extraction.
| Parameter |
Meaning |
pdf_dir |
Directory containing top-level *.pdf files with distinct filename stems, including when compared without letter case. |
extract_fn |
Callable accepting pdf_path and save_dir keyword arguments. The callable writes top-level PNGs into the supplied temporary directory. |
Returns None. After an entire PDF succeeds, Flowde publishes that PDF's PNGs
into the run's save_dir. The run record includes successful PDFs that produced
no images. Missing PNGs cause the whole source PDF to be extracted again on
resume; modified saved PNGs cause an error.
For the Paddle factory's crop names and settings, see
image extraction.
Classification
classify_imgs
classify_imgs(classify_fn: ClassificationFunction[LabelType], img_dir: Path, save_dir: Path, *, positive_classes: set[ClassificationLabel] | None = None, range_indices: tuple[int | None, int | None] | None = None, max_concurrent_jobs: int = 10, show_usage: bool = True, on_existing: ExistingRun = 'error') -> list[LabelType]
Classify PNG images in a directory and save their labels in a resumable run.
| Parameters: |
-
classify_fn
(ClassificationFunction[LabelType])
–
Function accepting an image Path as its first positional argument
and returning a str, int or bool label. The function returns the
label itself, not a dictionary or Pydantic model.
Built-in factories declare their settings;
custom functions must declare their settings with
model_function(). If the classifier exposes
a result_structure, returned labels are validated against that schema.
-
img_dir
(Path)
–
Directory containing the input *.png files. Subdirectories are not
searched. Image paths are sorted before applying range_indices.
Selected filename stems must be unique ignoring case across the run,
including images from earlier resumed calls.
-
save_dir
(Path)
–
Dedicated output directory, created if needed. Labels are saved in
classifications.json, run metadata in .flowde, and optional
image copies in positive_images. Input images and files referenced by
the classifier's declared settings must be outside this directory.
-
positive_classes
(set[ClassificationLabel] | None, default:
None
)
–
Labels whose images are copied into save_dir / "positive_images",
preserving filenames and leaving the original images unchanged.
For example, {1} copies images labelled 1. Defaults to None,
which saves labels without copying images. Must remain unchanged when
resuming the run.
-
range_indices
(tuple[int | None, int | None] | None, default:
None
)
–
A (start, stop) slice of the sorted image paths. start is included
and stop is excluded; (0, 10) selects up to the first ten images.
Either bound can be None, and negative indices follow Python slicing
rules. Defaults to None, which selects all matching images. An empty
selection raises an error.
-
max_concurrent_jobs
(int, default:
10
)
–
Maximum number of images handled concurrently by threads in a separate
worker process, by default 10. Must be a positive integer; 1 handles one
image at a time.
Custom functions must support concurrent calls when this exceeds 1.
Even at 1, worker mutations do not update objects in the caller; see
custom functions.
This setting can change when resuming a run.
-
show_usage
(bool, default:
True
)
–
Whether to display token usage and estimated costs reported by the
classifier. Defaults to True. False hides those figures while
retaining the progress display and saved usage reports. This value
can change when resuming a run.
-
on_existing
((error, resume, overwrite), default:
"error"
)
–
How to handle an existing run in save_dir. Defaults to "error".
"error": start in a new or empty directory; reject existing work.
"resume": require a saved classification run with matching classifier
settings and positive_classes. Reuse saved labels for unchanged
inputs and classify selected images without saved labels.
"overwrite": remove the previous run's tracked outputs and saved state,
then start a new run. Also works in a new or empty directory.
Unrelated files in an existing output directory cause an error.
Resume and overwrite require all internal records. Missing or invalid
records raise an error.
|
| Returns: |
-
list[LabelType]
–
One label per image selected by this call, in sorted image-path order.
Includes labels restored from a previous run. With range_indices, the
returned list covers only the selected slice; labels for earlier slices
remain saved in classifications.json.
|
| Raises: |
-
FileExistsError
–
If save_dir contains existing work and on_existing="error".
-
ValueError
–
If img_dir is missing or is not a directory, no PNGs are selected,
filename stems conflict ignoring case, inputs are inside save_dir,
or the saved run fails compatibility or integrity checks.
-
TypeError
–
If the classifier returns a label that is not a string, integer or
boolean, including None.
-
RuntimeError
–
If another Flowde call is already using the same save_dir.
|
Notes
classifications.json contains objects with img_path and label fields.
The JSON file accumulates labels across resumed calls and orders entries
by resolved input paths. Flowde saves settings in .flowde/run_metadata.state
and each image's label, usage and progress in .flowde/input_records.
Keep the whole .flowde directory with the outputs for resume and overwrite.
On resume, missing classifications.json or selected positive_images
copies are recreated from saved labels and unchanged input images without
calling the classifier again. Manually edited saved outputs raise an error.
See the classification guide for worked
examples.
Rotation
rotate_imgs
rotate_imgs(classify_fn: ClassificationFunction[RotationLabel], img_dir: Path, save_dir: Path, *, range_indices: tuple[int | None, int | None] | None = None, max_concurrent_jobs: int = 10, show_usage: bool = True, on_existing: ExistingRun = 'error') -> list[RotationLabel]
Rotate PNG images in a directory and save corrected copies in a resumable run.
| Parameters: |
-
classify_fn
(ClassificationFunction[RotationLabel])
–
Function accepting an image Path as its first positional argument
and returning a clockwise correction in degrees as an int: 0,
90, 180 or 270. The function returns the angle itself, not a
dictionary or Pydantic model, and must raise an exception if angle
prediction fails. Returning None raises an error.
The classifier must leave the input image unchanged. Flowde applies
the returned angle and saves the corrected copy. Built-in factories
declare their settings; custom functions must declare their settings with
model_function().
-
img_dir
(Path)
–
Directory containing the original *.png files. Subdirectories are not
searched. Image paths are sorted before applying range_indices.
Selected filename stems must be unique ignoring case across the run,
including images from earlier resumed calls.
-
save_dir
(Path)
–
Dedicated output directory, created if needed. Correction angles are
saved in rotations.json, corrected copies in rotated_images, and run
metadata in .flowde. Original images remain unchanged.
Input images and files referenced by the classifier's declared settings
must be outside this directory.
-
range_indices
(tuple[int | None, int | None] | None, default:
None
)
–
A (start, stop) slice of the sorted image paths. start is included
and stop is excluded; (0, 10) selects up to the first ten images.
Either bound can be None, and negative indices follow Python slicing
rules. Defaults to None, which selects all matching images. An empty
selection raises an error.
-
max_concurrent_jobs
(int, default:
10
)
–
Maximum number of images handled concurrently by threads in a separate
worker process, by default 10. Must be a positive integer; 1 handles one
image at a time.
Custom functions must support concurrent calls when this exceeds 1.
Even at 1, worker mutations do not update objects in the caller; see
custom functions.
This setting can change when resuming a run.
-
show_usage
(bool, default:
True
)
–
Whether to display token usage and estimated costs reported by the
classifier. Defaults to True. False hides those figures while
retaining the progress display and saved usage reports. This value
can change when resuming a run.
-
on_existing
((error, resume, overwrite), default:
"error"
)
–
How to handle an existing run in save_dir. Defaults to "error".
"error": start in a new or empty directory; reject existing work.
"resume": require a saved rotation run with matching classifier
settings. Reuse saved angles for unchanged inputs and predict angles
for selected images without saved angles.
"overwrite": remove the previous run's tracked outputs and saved state,
then start a new run. Also works in a new or empty directory.
Unrelated files in an existing output directory cause an error.
Resume and overwrite require all internal records. Missing or invalid
records raise an error.
|
| Returns: |
-
list[RotationLabel]
–
One clockwise correction angle per image selected by this call, in sorted
image-path order. Includes angles restored from a previous run. With
range_indices, the returned list covers only the selected slice;
angles for earlier slices remain saved in rotations.json.
|
| Raises: |
-
FileExistsError
–
If save_dir contains existing work and on_existing="error".
-
ValueError
–
If img_dir is missing or is not a directory, no PNGs are selected,
filename stems conflict ignoring case, an angle is unsupported, inputs
are inside save_dir, or the saved run fails compatibility or integrity
checks.
-
TypeError
–
If the classifier returns None instead of an angle.
-
RuntimeError
–
If another Flowde call is already using the same save_dir.
|
Notes
rotations.json contains objects with img_path and label fields, where
label is the clockwise correction angle. The JSON file accumulates angles
across resumed calls and orders entries by resolved input paths.
Flowde saves settings in .flowde/run_metadata.state and each image's angle,
usage and progress in .flowde/input_records. Keep the whole .flowde
directory with the outputs for resume and overwrite.
Every successfully processed image has a copy in rotated_images, including
images with a zero-degree correction. Copies retain their filenames and
expand to fit the rotated image without cropping.
On resume, missing rotations.json or selected rotated_images copies are
recreated from saved angles and unchanged original images without another
prediction. Corrected copies are always made from the original inputs, so
resuming does not rotate an already-corrected copy again. Edited previously
processed input images or saved outputs cause an error.
See the rotation guide for worked examples.
Parsing
parse_imgs
parse_imgs(parse_fn: ParsingFunction, img_dir: Path, save_dir: Path, nodes_dir: Path | None = None, labels_dir: Path | None = None, additional_texts_dir: Path | None = None, flow_dir: Path | None = None, range_indices: tuple[int | None, int | None] | None = None, max_concurrent_jobs: int = 10, img_extensions: set[str] | None = None, *, show_usage: bool = True, on_existing: ExistingRun = 'error') -> list[BaseModel]
Parse images in a directory into JSON files, optionally using saved context.
| Parameters: |
-
parse_fn
(ParsingFunction)
–
Function accepting an img_path keyword argument and, when context is
supplied, a partial_flowchart keyword argument containing a Pydantic
model. img_path is a Path to the image being parsed; the partial
model contains previously parsed parts for that same image. Without
saved context, Flowde passes only img_path.
The function must return an instance of the Pydantic class exposed
as parse_fn.result_structure, not a dictionary, JSON string or None.
The class can describe standard flowchart parts or a custom output
format. The parser must raise an exception if parsing fails. Flowde
saves each returned model as a same-stem JSON file in save_dir.
Built-in factories declare their settings;
custom functions must declare their settings and result class with
model_function().
-
img_dir
(Path)
–
Directory containing input images. Subdirectories are not searched.
Matching paths are sorted before applying range_indices. Filename stems
must be unique across all matching images, including different extensions,
because each stem determines an output JSON filename.
Selected stems must also remain unique ignoring case across the run,
including images from earlier resumed calls.
-
save_dir
(Path)
–
Dedicated output directory, created if needed. Each image produces a JSON
file with the same stem, such as diagram.png producing diagram.json.
Run metadata is saved in .flowde. Input images, partial JSONs
and files referenced by the parser's declared settings must be outside
this directory.
-
nodes_dir
(Path | None, default:
None
)
–
Directory of node-text JSONs supplied as context alongside the images.
Each JSON contains nodes with node_number and text fields.
Defaults to None, meaning no saved context is supplied. Required if
labels_dir, additional_texts_dir or flow_dir is provided.
-
labels_dir
(Path | None, default:
None
)
–
Directory of label JSONs to add to the context from nodes_dir.
Each JSON contains nodes with node_number and labels fields.
Defaults to None, meaning no saved labels are included in the context.
-
additional_texts_dir
(Path | None, default:
None
)
–
Directory of additional-text JSONs to add to the context from nodes_dir.
Each JSON contains an additional_texts list. Defaults to None, meaning
no saved additional text is included in the context.
-
flow_dir
(Path | None, default:
None
)
–
Directory of connection JSONs to add to the context from nodes_dir.
Each JSON contains nodes with node_number and points_to fields.
Defaults to None, meaning no saved connections are included in the
context.
-
range_indices
(tuple[int | None, int | None] | None, default:
None
)
–
A (start, stop) slice of the sorted image paths. start is included
and stop is excluded; (0, 10) selects up to the first ten images.
Either bound can be None, and negative indices follow Python slicing
rules. Defaults to None, which selects all matching images. The same
slice is applied to supplied context files after checking their filenames
against all matching images. An empty selection raises an error.
-
max_concurrent_jobs
(int, default:
10
)
–
Maximum number of images handled concurrently by threads in a separate
worker process, by default 10. Must be a positive integer; 1 handles one
image at a time.
Custom functions must support concurrent calls when this exceeds 1.
Even at 1, worker mutations do not update objects in the caller; see
custom functions.
This setting can change when resuming a run.
-
img_extensions
(set[str] | None, default:
None
)
–
Image extensions to select, without leading dots, such as
{"png", "jpg", "webp"}. Defaults to None, which selects PNG files.
The supplied parser must support the selected image formats.
-
show_usage
(bool, default:
True
)
–
Whether to display token usage and estimated costs reported by the
parser. Defaults to True. False hides those figures while retaining
the progress display and saved usage reports. This value can change
when resuming a run.
-
on_existing
((error, resume, overwrite), default:
"error"
)
–
How to handle an existing run in save_dir. Defaults to "error".
"error": start in a new or empty directory; reject existing work.
"resume": require a saved parsing run with matching parser settings.
Restore saved results for unchanged images and assembled partial
context, and parse selected images without saved results.
"overwrite": remove the previous run's tracked outputs and saved state,
then start a new run. Also works in a new or empty directory.
Unrelated files in an existing output directory cause an error.
Resume and overwrite require all internal records. Missing or invalid
records raise an error.
|
| Returns: |
-
list[BaseModel]
–
One Pydantic result per image selected by this call, in sorted image-path
order. Each result uses parse_fn.result_structure. Includes results
restored from a previous run. With range_indices, the returned list
covers only the selected slice; JSON files for earlier slices remain
saved in save_dir.
|
| Raises: |
-
FileExistsError
–
If save_dir contains existing work and on_existing="error".
-
ValueError
–
If img_dir is invalid, no images are selected, image stems conflict
(including case-only differences across the run), context files do not
match the images or required part schemas,
node numbers cannot be joined, inputs are inside save_dir, or the saved
run fails compatibility or integrity checks.
-
TypeError
–
If parse_fn.result_structure is not a Pydantic class or the parser
returns a value that is not an instance of that class, including None.
-
RuntimeError
–
If another Flowde call is already using the same save_dir.
|
Notes
Each supplied context directory must contain exactly one top-level *.json
file per matching input image, with the same stem and no extra JSON files.
Filename and file-count checks cover all matching images before slicing,
including images outside range_indices.
Context files contain the fields for their individual parts, without the
benchmark ground truth's options wrapper. Selected node-text files must
contain at least one node, numbered consecutively from 1. Selected label
and flow files must contain the same node numbers as the node-text file.
The parts are joined into the partial_flowchart passed to the parser.
Context files supply input to the parser; they are not automatically merged
into its output. The output fields are defined by parse_fn.result_structure.
For built-in parsers, the factory's parts_to_parse or result_structure
argument chooses those fields.
Flowde saves settings in .flowde/run_metadata.state and each image's result,
usage and progress in .flowde/input_records/<stem>.state. Progress updates
rewrite only that image's record, without rewriting other images' results
or the run metadata. Keep the whole .flowde directory with the outputs.
Flowde records each result before writing its output JSON. On resume, missing
output JSONs are recreated from these records without another parsing request.
Edited saved outputs or changes to previously parsed input images or their
assembled partial context cause an error. Missing or invalid internal records
cause an error even when the output JSONs still exist.
See the parsing guide for full, partial and
custom-schema examples.