Document Loader

processing.document-loader Processing v0.1.0

Loads text from a source so it can be chunked, embedded, and stored — the on-ramp for RAG ingestion. Fetch a web page or document by URL (optionally reducing HTML to readable text) or read a text field already on the input, and emit { text, chars, source } that chains straight into the Text Splitter node.

The Document Loader step on the Studio canvas
The Document Loader step as it appears on the Studio canvas — input pins on the left, output ports on the right.

Finding it in the library

Search the builder's node library for Document Loader (it lives under Processing). A single click opens the in-editor docs panel shown here — description, ports, and every property, without leaving the canvas. Double-click (or drag) to add it to the workflow.

Document Loader in the node library, with the in-editor docs panel open
The library entry and the in-editor docs panel for Document Loader — the same reference this page is generated from.

Wired up in the builder

Document Loader in a real, runnable flow — captured live from the Studio editor, exactly as it looks on your canvas. This is the same workflow used for the example input & output below.

Document Loader wired into a runnable workflow in the Studio builder
Document Loader wired into a runnable flow — input on the left, output on the right.

How it’s configured

The node’s settings as the builder shows them — every field laid out with real values. In the Studio these are edited on the node: click the chevron on the divider under its ports to open them.

The Document Loader node's settings in the Studio builder
The settings for Document Loader, showing the values from the flow above.

Ports

Ports are the node’s contract with its neighbours. In the editor a port label renders bold when wired and italic when optional; ports accept attachment carriers rather than data wires.

DirectionPortLabelWhat flows through it
InputinputInput
OutputoutputText

How data flows through it

Document Loader consumes the content of the incoming envelope — when it is fed directly by a trigger, the trigger’s wrapper is unwrapped at the node boundary so the node sees the actual data, not the metadata shell. Its output becomes the payload for the next node, while the envelope (trace ids, correlation, binary refs) rides along untouched. In the Runs view you always see the whole envelope for both sides of this node.

Expressions in the config

String-typed properties accept {{ }} expressions evaluated against the incoming item at run time — e.g. {{ $json.customer.email }}. On this node that’s url. JSON- and code-typed fields never interpolate — they are passed through literally.

Build it with AI

Every node in this reference is reachable through Flowdrome’s AI Copilot and the MCP tools — say what you want, and the graph surgery happens server-side. Node types resolve fuzzily, so the catalog label (Document Loader) works as well as the exact type id (processing.document-loader).

In the Copilot panel (or any connected AI):

add a document loader node after the trigger

As a step in a create_chain_workflow call:

{"type":"Document Loader","config":{}}
Raw MCP call — add this node to a workflow with add_node
curl -s -X POST http://localhost:4800/mcp -H "content-type: application/json" -d '{ "jsonrpc": "2.0", "id": "1", "method": "tools/call", "params": { "name": "add_node", "arguments": { "workflowId": "<id>", "type": "Document Loader" } } }'

Example input & output

Captured from a real test run of the workflow above — this is what the Runs view shows after pressing Test workflow.

Input — what the node received

The Document Loader node's input envelope in the run data viewer
The input envelope in the Runs view — Flowdrome always shows the whole envelope, with the payload inside body.

Output — what the node produced

The Document Loader node's output envelope in the run data viewer
The output envelope after the step ran.

Property reference

Every setting, with its type and default — the same fields shown configured above.

PropertyTypeDefaultDescription
Source
source
select "url" Where the text comes from: url (fetch a web page or document over http(s)) or field (read a text field already on the input).
fieldurl
URL
url
string "" The http(s) URL to fetch (GET).
Shown when (source ?? "url") === "url"
Text field
field
field "" Dot-path to the text field on the input. Blank = the whole payload (stringified).
Shown when (source ?? "url") === "field"
Strip HTML
stripHtml
boolean true For URLs, reduce HTML markup to readable text before emitting.
Shown when (source ?? "url") === "url"
Timeout (ms)
timeoutMs
int 30000 Abort the fetch after this many milliseconds.
Shown when (source ?? "url") === "url"

Related nodes

The rest of the Processing group — the same folder you’d scan in the editor’s library.

This page is generated from the node registry by gen-node-docs.mjs on every site build — ports, properties, defaults and visibility rules cannot drift from the code. The screenshots and example data are captured from a live Flowdrome by npm run shots:nodes and npm run gen:examples.