Data types

Classification schemas

Import these types from flowde.classify_fns.classify_types.

Type Output
Classification[LabelType] Pydantic model with a single label field. Specialise the generic with a type or a Literal of allowed values.
BinaryClassification label is integer 0 or 1.
ConsortClassification Binary schema for the supplied CONSORT classification prompt: 0 for other images, 1 for CONSORT diagrams.
RotationClassification label is 0, 90, 180 or 270: the clockwise correction to apply.

Classification schemas use strict validation. For example, the JSON string "1" is not the same as the integer 1 required by BinaryClassification. Model factories return the validated label from the response.

from typing import Literal

from flowde.classify_fns.classify_types import Classification

ImageType = Classification[Literal["flowchart", "table", "other"]]
answer = ImageType.model_validate({"label": "flowchart"})
print(answer.label)

Parsing schemas

Import ParseType, Node, Flowchart and the schema-building helpers from flowde.parsing_fns.parsing_types.

ParseType value JSON fields requested
"node_text" nodes, with node_number and text for every node
"labels" nodes, with node_number and labels for every node
"flow" nodes, with node_number and points_to for every node
"additional_texts" Top-level additional_texts list

Node has node_number: int, text: str, labels: list[str] and points_to: list[int]. Flowchart has nodes: list[Node] and additional_texts: list[str]. The default parsing factories create a schema for the selected parts using build_partial_flowchart_schema() .

Inference writes one JSON object per image. Ground truth uses an options wrapper and separates the four parts into different directories. See the parsing benchmark format.

Pydantic results support model_dump() for a Python dictionary, model_dump_json(indent=2) for JSON text, and model_json_schema() for the schema. See parsing for complete examples.

Few-shot example

VisionFewShotExample dataclass

VisionFewShotExample(*, img_path: Path, expected_output_path: Path, partial_flowchart: BaseModel | None = None)

Describe one example image and its expected JSON answer for a model prompt.

Parameters:
  • img_path (Path) –

    Path to an existing example image. The image is included in the prompt as a worked example, before the target image being classified or parsed. The image format must be supported by the selected provider.

  • expected_output_path (Path) –

    Path to an existing UTF-8 JSON file containing the desired response for the example image. For binary classification, the file could contain {"label": 1}. For parsing, the JSON should match the requested parts or custom response schema, without a benchmark ground truth's options wrapper. Schema compatibility is the caller's responsibility; this class does not validate the example answer against the model's response schema.

  • partial_flowchart (BaseModel | None, default: None ) –

    Previously parsed data supplied as input context for this example image, such as node numbers and text when demonstrating label parsing. Defaults to None, meaning no partial context accompanies the example. This context belongs to img_path, not to the target image that the returned classifier or parser will later process.

Raises:
  • ValueError –

    If img_path or expected_output_path does not point to an existing file when the example is created.

Notes

Constructor arguments must be passed by name. The dataclass is frozen, so its fields cannot be reassigned after construction. The referenced files and the supplied Pydantic model are not made immutable.

Creating an example checks that both paths point to files; it does not decode the image, read the expected JSON or make an API request. expected_output_text() reads and formats the expected JSON when a request is prepared. File contents are read from disk rather than cached in this object.

You can pass a list of these instances as few_shot_examples to an OpenAI or Gemini classification or parsing factory. Each request includes the worked examples in list order, followed by the target image. Example answers demonstrate the desired response; they are not predictions for the target image.

When a factory's callable is used in a saved Flowde run, the example image, expected-response file and partial context contribute to the settings checked on resume. Changes to those files or that context require a new run or explicit overwrite.

Examples:

Given an existing flowchart.png and a flowchart.json containing {"label": 1}, create a classification example:

from pathlib import Path

from flowde.utils import VisionFewShotExample

example = VisionFewShotExample(
    img_path=Path("data/examples/flowchart.png"),
    expected_output_path=Path("data/examples/flowchart.json"),
)

The factory's few_shot_examples argument can then receive [example].

expected_output_text

expected_output_text() -> str

Read the expected-response file and return its contents as formatted JSON.

Returns:
  • str –

    JSON text with two-space indentation. Non-ASCII characters are preserved rather than escaped. The file is read afresh on every call; no response-schema validation is performed.

Raises:
  • OSError –

    If expected_output_path cannot be opened or read, including when the file has been removed after the example was created.

  • UnicodeDecodeError –

    If the expected-response file is not valid UTF-8 text.

  • JSONDecodeError –

    If the expected-response file does not contain valid JSON.

See few-shot examples for classification and parsing examples using model factories.

Token prices

TokenPrices dataclass

TokenPrices(*, input: float, output: float, cached_input: float | None = None, cache_write: float | None = None, long_context_threshold: int | None = None)

Store token prices for estimating the cost of a model request.

Parameters:
  • input (float) –

    Price in USD per million uncached input tokens.

  • output (float) –

    Price in USD per million output tokens.

  • cached_input (float | None, default: None ) –

    Price in USD per million input tokens read from a cache. This rate replaces the ordinary input rate for those tokens. Defaults to None, meaning the cache-read rate is unknown. A value of 0 means cache reads are free. An unknown rate prevents a cost estimate only when the request includes cache-read tokens.

  • cache_write (float | None, default: None ) –

    Price in USD per million input tokens written to a cache. This rate replaces the ordinary input rate for those tokens. Defaults to None, meaning the cache-write rate is unknown. A value of 0 means cache writes are free. An unknown rate prevents a cost estimate only when the request includes cache-write tokens.

  • long_context_threshold (int | None, default: None ) –

    Input-token count above which higher rates apply. Defaults to None, meaning the same rates apply at every request length. When a request's total input-token count is strictly greater than this threshold, all input rates, including cache reads and writes, are multiplied by 2, and the output rate is multiplied by 1.5. The higher rates apply to the entire request, not just tokens beyond the threshold. These multipliers are fixed; this option is suitable only for prices that follow that rule.

Notes

Constructor arguments must be passed by name. The dataclass is frozen, so its fields cannot be reassigned after construction. Rates and the threshold are stored as supplied, without numeric-range validation or a provider price lookup.

You can pass an instance as token_prices to an OpenAI or Gemini model factory to override Flowde's bundled prices for cost reporting. The override changes estimates, not provider billing or model requests. estimate() can also calculate a cost directly from token counts. Creating a TokenPrices instance or calculating a cost makes no API request.

Examples:

Define illustrative rates of $2 per million uncached input tokens, $8 per million output tokens and $0.50 per million cached input tokens:

from flowde.pricing import TokenPrices

prices = TokenPrices(input=2.0, output=8.0, cached_input=0.5)

These are example rates, not prices for a particular model.

estimate

estimate(input_tokens: int | None, output_tokens: int | None, *, cached_tokens: int = 0, cache_write_tokens: int = 0) -> float | None

Estimate one request's cost from its input, output and cache token counts.

Parameters:
  • input_tokens (int | None) –

    Total input tokens, including the tokens counted in cached_tokens and cache_write_tokens. None means the count is unknown and prevents a cost estimate. This total determines whether the long-context threshold is exceeded.

  • output_tokens (int | None) –

    Total output tokens to price at the output rate. Include any reasoning or thinking tokens charged as output. None means the count is unknown and prevents a cost estimate.

  • cached_tokens (int, default: 0 ) –

    Input tokens read from a cache, priced at cached_input instead of input. Defaults to 0. These tokens must already be included in input_tokens and must not also count as cache-write tokens.

  • cache_write_tokens (int, default: 0 ) –

    Input tokens written to a cache, priced at cache_write instead of input. Defaults to 0. These tokens must already be included in input_tokens and must not also count as cache-read tokens.

Returns:
  • float | None –

    Estimated cost in USD. Returns None if either total token count is unknown, a nonzero cache count has no corresponding price, or the two cache counts together exceed input_tokens. Known zero counts or zero prices can produce a cost of 0.0.

Notes

Uncached input tokens are calculated as input_tokens - cached_tokens - cache_write_tokens. Each input group is multiplied by its own per-million-token rate; output tokens are multiplied by the output rate. The sum is divided by 1_000_000.

When input_tokens exceeds a configured long_context_threshold, the combined input cost is doubled and the output cost is multiplied by 1.5 before summing. At exactly the threshold, the original rates apply. Output tokens do not contribute to the threshold comparison.

Counts are expected to be nonnegative integers. Apart from the checks described under Returns, counts and prices are used as supplied. The method does not validate every numeric range, contact a provider, round the estimate or include charges unrelated to these token counts.

Examples:

Calculate a cost using illustrative rates and 1,000 input tokens, of which 400 were read from a cache, plus 200 output tokens:

from flowde.pricing import TokenPrices

prices = TokenPrices(input=2.0, output=8.0, cached_input=0.5)
cost = prices.estimate(
    input_tokens=1_000,
    output_tokens=200,
    cached_tokens=400,
)

The estimated cost is $0.003: $0.0012 for the 600 uncached input tokens, $0.0002 for the 400 cached input tokens and $0.0016 for the 200 output tokens.

See tokens and costs for usage reporting and custom-price examples.

Usage reports

RequestUsage dataclass

RequestUsage(*, model: str, provider: str, input_tokens: int | None = None, output_tokens: int | None = None, total_tokens: int | None = None, cost: float | None = None)

Usage from one contributing operation; unknown values remain None.

Create a report with model and provider, plus any known input_tokens, output_tokens, total_tokens and cost. Unknown values default to None. For a custom function, supply total_tokens explicitly if you want the total to appear in Flowde's summary; the report object does not calculate that field.

report_usage

report_usage(usage: RequestUsage) -> None

Report a contributing operation without changing the callable's return value.

Outside a tracked run this does nothing. Custom functions may call this for each model response or separately known charge; reporting is never required. Costs are USD and token counts must describe this operation, not a running sum.

Records a RequestUsage report for the currently running item. Custom classifiers and parsers can report more than one request per image. Calling report_usage() outside a managed usage context has no effect. See custom usage reporting.

Function contracts

These protocols describe the callable interfaces that Flowde's processing functions must follow. Custom functions do not need to inherit from a protocol.

Pipeline stage Function contract Result for one input file
Extraction ExtractImgsFunction Write PNG files for one PDF into the supplied directory; return None.
Classification and rotation ClassificationFunction Return one label, or a clockwise correction angle for rotation.
Parsing ParsingFunction Return a Pydantic model matching the callable's result_structure class.

The contracts below describe the arguments, return values and requirements for each callable. See the custom-function guide for working examples and instructions on declaring settings for saved runs.

ExtractImgsFunction

Bases: Protocol

Describe a callable that extracts PNG images from one PDF.

Pass an extractor with this interface as extract_fn to extract_imgs(). Flowde supplies one PDF and a temporary output directory per call. The extractor writes the PNG files; Flowde validates the files and publishes the completed PDF's images in the run's output directory.

Notes

This protocol describes an interface; it does not implement extraction or perform runtime validation. A function or callable object can satisfy the interface without inheriting from ExtractImgsFunction.

Custom extractors must also declare their settings with model_function(). The resulting run_settings attribute lets Flowde check compatibility when resuming a saved run. This requirement applies in addition to the call signature defined here. Built-in extractor factories declare their settings automatically.

__call__

__call__(*, pdf_path: Path, save_dir: Path) -> None

Extract images from one PDF into the supplied temporary directory.

Parameters:
  • pdf_path (Path) –

    Path to the input PDF. Flowde supplies this argument by keyword. The extractor must leave the input PDF unchanged.

  • save_dir (Path) –

    Existing, empty temporary directory for this PDF. Flowde supplies this argument by keyword. This directory is separate from the final save_dir passed to extract_imgs().

    Save only valid PNG files with the lowercase .png extension directly inside this directory. Do not create subdirectories, symbolic links or other files. Filenames must be unique across all images in the run, including when compared without letter case. Including the PDF's filename stem in each PNG's name helps avoid collisions between PDFs.

Returns:
  • None –

    The saved PNG files are the extraction results. Finish writing and close every output file before returning. Leaving save_dir empty is valid when the PDF contains no images to extract.

Notes

Raise an exception if extraction fails. Flowde publishes a PDF's PNG files only after the extractor returns successfully and the PNGs pass validation. An extractor must not leave background work writing files after returning.

ClassificationFunction

Bases: Protocol[LabelType_co]

Describe a callable that returns one classification label for one image.

Pass a classifier with this interface as classify_fn to classify_imgs(). The generic label type describes the callable's return value: a str, int or bool. For example, ClassificationFunction[str] returns string labels.

Rotation uses the same interface. A ClassificationFunction[RotationLabel] passed to rotate_imgs() must return the clockwise correction angle as an integer: 0, 90, 180 or 270. The classifier predicts the angle; Flowde applies the rotation and saves corrected image copies.

Notes

This protocol describes an interface; it does not implement classification or perform runtime validation. A function or callable object can satisfy the interface without inheriting from ClassificationFunction.

Custom classifiers must also declare their settings with model_function(). The resulting run_settings attribute lets Flowde check compatibility when resuming a saved run. This requirement applies in addition to the call signature defined here. Built-in classifier factories declare their settings automatically.

A classifier can optionally expose a result_structure class with a label field to restrict the permitted labels. For example, BinaryClassification permits only the integers 0 and 1. The classification pipeline checks each returned label against the supplied schema. Without a schema, the classification pipeline accepts any str, int or bool label. The rotation pipeline always checks for one of the four correction angles.

__call__

__call__(img_path: Path) -> LabelType_co

Classify one image and return its label or clockwise correction angle.

Parameters:
  • img_path (Path) –

    Path to the target image. Flowde supplies the path as the first positional argument. The classifier must support the image's format and leave the source image unchanged.

Returns:
  • LabelType_co –

    One str, int or bool label matching the callable's declared generic label type. Return the label itself, not a dictionary, JSON string containing an object, or Pydantic model. A rotation classifier must return the integer clockwise correction angle 0, 90, 180 or 270; 0 means no correction is needed.

Notes

Raise an exception if classification or angle prediction fails. Returning None causes a pipeline error and does not mark the image as successfully processed.

ParsingFunction

Bases: Protocol

Describe a callable that parses one image into a Pydantic result.

Attributes:
  • result_structure (type[BaseModel]) –

    Pydantic class describing the parser's output. Every successful call must return an instance of this class. The attribute holds the class, not an already parsed result. The class can describe a complete flowchart, selected flowchart parts or a custom output format.

Notes

This protocol describes an interface; it does not implement image parsing. A function or callable object can satisfy the interface without inheriting from ParsingFunction. The callable accepts img_path and optional partial_flowchart arguments, as described by __call__().

The OpenAI and Gemini parsing factories return callables with this interface. A returned parser processes one image per call and returns its Pydantic result. You can pass the parser to parse_imgs() to process a directory and save one JSON file per image.

A custom parser used by Flowde must also expose its declared run_settings. You can attach these settings and the required result_structure class with model_function(). These settings allow saved-run compatibility checks; run_settings is a pipeline requirement beyond the attributes declared by this protocol. Built-in factories attach these settings automatically.

The protocol itself performs no runtime validation. The directory pipeline checks that result_structure is a Pydantic class and that each returned result is an instance of that class. A dictionary, JSON string or None is not an accepted parsing result.

__call__

__call__(img_path: Path, partial_flowchart: BaseModel | None = None) -> BaseModel

Parse one image, optionally using previously parsed data as context.

Parameters:
  • img_path (Path) –

    Path to the target image. The parser implementation must support the image's format. Flowde supplies this argument by keyword.

  • partial_flowchart (BaseModel | None, default: None ) –

    Previously parsed data for the same image, such as node numbers and text when requesting labels. Defaults to None, meaning no earlier results are supplied. Flowde supplies this argument by keyword when context files are provided to parse_imgs().

Returns:
  • BaseModel –

    Parsed result for the target image, as an instance of the callable's result_structure class. The output schema determines which fields the result contains. Supplying partial context does not require those context fields to appear in the result.

Notes

This method defines the call signature; a parser implementation performs the actual work. The implementation can raise an exception when parsing fails. Returning None does not mark an image as successfully parsed.