Set up the default functions
Flowde handles directories, saved results, parallel workers and benchmarking. The function supplied to each pipeline stage does the actual extraction, classification or parsing. You can use a default helper or bring your own function.
Set up image extraction
Install the PaddleOCR dependencies, then create an extractor:
from flowde.extract_fns.paddle_layout_detect_extraction import (
make_paddle_layout_extract_fn,
)
extract_fn = make_paddle_layout_extract_fn(device="cpu", cpu_threads=1)
The layout model is loaded when extraction starts. The first run may download
model files. Use device="gpu" only after completing the
GPU setup.
See image extraction for the function that processes your PDF directory.
Set up OpenAI
Install the OpenAI extra. In the directory
where you will run your script or notebook, create a .env file containing:
OPENAI_API_KEY=your-openai-api-key
Load that file explicitly at the start of your example:
from pathlib import Path
from dotenv import load_dotenv
load_dotenv(Path(".env"))
Flowde's helpers also load environment settings. Explicitly loading the file above makes its location clear when your script and working directory differ.
With your API key loaded, you can create functions that use OpenAI's Responses API to classify, rotate or parse images:
| Stage | Create the function with |
|---|---|
| Classification | make_openai_classify_fn(...) |
| Rotation | make_openai_classify_fn( |
| Parsing | make_openai_parse_fn(...) |
For example, create a classifier:
from flowde.classify_fns.openai_classify_fn import make_openai_classify_fn
classify_fn = make_openai_classify_fn(
input_text="Return 1 if the image is a flowchart, otherwise return 0.",
model="gpt-5.6-luna",
effort="medium",
)
Creating the function does not classify any images. Pass classify_fn to
classify_imgs() to run
the requests. Model requests are
billable through your provider account.
Use Azure OpenAI
Install the OpenAI extra. In the directory
where you will run your script or notebook, create a .env file containing:
AZURE_API_KEY=your-azure-api-key
AZURE_API_BASE=https://your-resource.openai.azure.com/openai/v1/
Load the .env file as shown in Set up OpenAI, then use
from_azure=True to send requests through Azure:
| Stage | Create the function with |
|---|---|
| Classification | make_openai_classify_fn( |
| Rotation | make_openai_classify_fn( |
| Parsing | make_openai_parse_fn( |
For example, create a classifier:
from flowde.classify_fns.openai_classify_fn import make_openai_classify_fn
classify_fn = make_openai_classify_fn(
input_text="Return 1 if the image is a flowchart, otherwise return 0.",
model="your-deployment-name",
effort="medium",
from_azure=True,
)
Use Gemini
Install the Gemini extra. In the directory
where you will run your script or notebook, create a .env file containing:
GOOGLE_GENAI_API_KEY=your-gemini-api-key
Load the .env file as shown in Set up OpenAI, then create
functions that use the Gemini API to classify, rotate or parse images:
| Stage | Create the function with |
|---|---|
| Classification | make_gemini_classify_fn(...) |
| Rotation | make_gemini_classify_fn( |
| Parsing | make_gemini_parse_fn(...) |
For example, create a classifier:
from flowde.classify_fns.gemini_classify_fn import make_gemini_classify_fn
classify_fn = make_gemini_classify_fn(
input_text="Return 1 if the image is a flowchart, otherwise return 0.",
model="gemini-3.1-flash-lite",
effort="high",
)
Next step
Extract images, or start with classification if you already have PNGs. The API reference lists the helper parameters.