Set up the default functions

Flowde handles directories, saved results, parallel workers and benchmarking. The function supplied to each pipeline stage does the actual extraction, classification or parsing. You can use a default helper or bring your own function.

Set up image extraction

Install the PaddleOCR dependencies, then create an extractor:

from flowde.extract_fns.paddle_layout_detect_extraction import (
    make_paddle_layout_extract_fn,
)

extract_fn = make_paddle_layout_extract_fn(device="cpu", cpu_threads=1)

The layout model is loaded when extraction starts. The first run may download model files. Use device="gpu" only after completing the GPU setup.

See image extraction for the function that processes your PDF directory.

Set up OpenAI

Install the OpenAI extra. In the directory where you will run your script or notebook, create a .env file containing:

OPENAI_API_KEY=your-openai-api-key

Load that file explicitly at the start of your example:

from pathlib import Path

from dotenv import load_dotenv

load_dotenv(Path(".env"))

Flowde's helpers also load environment settings. Explicitly loading the file above makes its location clear when your script and working directory differ.

With your API key loaded, you can create functions that use OpenAI's Responses API to classify, rotate or parse images:

Stage Create the function with
Classification make_openai_classify_fn(...)
Rotation make_openai_classify_fn(
    ...,
    result_structure=RotationClassification,
)
Parsing make_openai_parse_fn(...)

For example, create a classifier:

from flowde.classify_fns.openai_classify_fn import make_openai_classify_fn

classify_fn = make_openai_classify_fn(
    input_text="Return 1 if the image is a flowchart, otherwise return 0.",
    model="gpt-5.6-luna",
    effort="medium",
)

Creating the function does not classify any images. Pass classify_fn to classify_imgs() to run the requests. Model requests are billable through your provider account.

Use Azure OpenAI

Install the OpenAI extra. In the directory where you will run your script or notebook, create a .env file containing:

AZURE_API_KEY=your-azure-api-key
AZURE_API_BASE=https://your-resource.openai.azure.com/openai/v1/

Load the .env file as shown in Set up OpenAI, then use from_azure=True to send requests through Azure:

Stage Create the function with
Classification make_openai_classify_fn(
    ...,
    from_azure=True,
)
Rotation make_openai_classify_fn(
    ...,
    result_structure=RotationClassification,
    from_azure=True,
)
Parsing make_openai_parse_fn(
    ...,
    from_azure=True,
)

For example, create a classifier:

from flowde.classify_fns.openai_classify_fn import make_openai_classify_fn

classify_fn = make_openai_classify_fn(
    input_text="Return 1 if the image is a flowchart, otherwise return 0.",
    model="your-deployment-name",
    effort="medium",
    from_azure=True,
)

Use Gemini

Install the Gemini extra. In the directory where you will run your script or notebook, create a .env file containing:

GOOGLE_GENAI_API_KEY=your-gemini-api-key

Load the .env file as shown in Set up OpenAI, then create functions that use the Gemini API to classify, rotate or parse images:

Stage Create the function with
Classification make_gemini_classify_fn(...)
Rotation make_gemini_classify_fn(
    ...,
    result_structure=RotationClassification,
)
Parsing make_gemini_parse_fn(...)

For example, create a classifier:

from flowde.classify_fns.gemini_classify_fn import make_gemini_classify_fn

classify_fn = make_gemini_classify_fn(
    input_text="Return 1 if the image is a flowchart, otherwise return 0.",
    model="gemini-3.1-flash-lite",
    effort="high",
)

Next step

Extract images, or start with classification if you already have PNGs. The API reference lists the helper parameters.