Overview
Pipechat is a visual data-pipeline workspace. You describe the dataset you want, assemble the work as connected blocks, and run the generated project on managed compute — now or on a schedule. The graph, its code, its data, and its execution history stay connected instead of living in separate tools.
Most work follows the same four stages:
- CollectPoint a pipeline at a site. Pipechat samples real pages and writes structural extraction rules for it.
- StructureTitles, dates, bodies, sections, attachments, and hyperlinks come back as consistent fields.
- EnrichPython and LLM blocks clean, join, classify, and derive whatever the raw records are missing.
- DeliverWrite to managed storage or your own database, or expose the result as an MCP server.
Rules are written once, then executed. The model reads sample pages at build time and produces extraction rules. Every run after that executes those rules directly — there is no model in the collection path, so the same fields come back each time.
Core concepts
- Workspace
- The project boundary. A workspace groups related pipelines, managed storage resources, MCP servers, and multi-pipeline orchestrations.
- Pipeline
- A directed graph that collects, transforms, enriches, and writes data. Every pipeline belongs to one workspace.
- Block
- One executable step, with configuration, generated source files, inputs, outputs, and connections to other blocks.
- Run
- One execution of a pipeline. It records the trigger, timing, logs, status, and an input/output snapshot for each block.
- Orchestration
- An ordered chain of pipelines. Each step begins only when the pipeline before it succeeds.
A connection is both visual and executable: it determines which upstream output becomes the next block's input. Blocks run in connection order — records leave a source, move through transforms, and arrive at a destination.
Collecting web data
An External Sources block collects listing pages and individual content from websites, APIs, and public feeds. Hosted web collection fetches current page content at run time.
When Pipechat builds a source, it captures a small representative sample of real pages and derives extraction rules from them. Those rules are validated against the captures before they are saved, so a rule that would collect nothing in production fails at build time instead of quietly returning empty rows.
URLs are normalised on the way in: tracking parameters are stripped, hosts are lowercased, http is upgraded to https, and fragments are dropped — so the same page doesn't arrive as several distinct records.
Source Library
The Source Library holds reusable, maintained sources. A source records the address, publisher, availability, extraction rules, and generated files needed to collect listings and individual pages. Availability tells you whether a source is ready to collect; connection state tells you whether it's already attached to the current pipeline. The definition is reusable, and each run still fetches current content.
Block types
- External Sources
- Collects listing pages and individual content from websites, APIs, and public feeds.
- Custom Code
- Runs a specialised JavaScript or Python transformation when a purpose-built block does not fit.
- LLM Transform
- Extracts, classifies, summarises, or reshapes records using your prompt and a linked model-provider key.
- Embeddings
- Converts text into vectors for semantic search, clustering, matching, or retrieval.
- Database Connector
- Reads from or writes to a database, including a relation created from the workspace's Storage tab.
Chat and code files
The chat panel can add, update, connect, move, and remove blocks. Describe the result and its constraints — for example, “collect press releases, extract company and funding amount, then store unique records” — and the agent updates both the graph and its project files.
The Code Files tab exposes the generated project behind the canvas. Each block maps to readable files, so you can inspect how it works, edit precise behaviour, and verify the agent's changes.
Be specific. Input shape, required output fields, deduplication rules, and failure behaviour give the agent a far better contract than a broad request such as “clean the data”.
Runs and debugging
Running a pipeline materialises the current project, starts a managed execution, and records the outcome. Run History combines manual, scheduled, and orchestration-triggered executions in one place, each row showing status, trigger, start time, duration, and output summary.
Lifecycle
- GeneratePipechat builds the executable project from the current graph and code files.
- ExecuteBlocks run in dependency order with the configured limits and linked credentials.
- CaptureInputs, outputs, logs, timing, and errors are saved as each block finishes.
- PublishThe run is marked succeeded only after every required block completes.
Debugging a failure
Start at the first failed block, not the final destination. Read its error and log, inspect its input snapshot, then compare that shape with what its code expects. Common causes are a changed source page, a missing field, an unlinked key, an invalid response from an external service, or a compute timeout.
After a fix, run the pipeline again. Pipechat preserves the failed attempt and creates a new run, giving you a before-and-after trail instead of overwriting the evidence.
Schedules and orchestration
The Scheduling tab creates recurring triggers for one pipeline. Enter a standard five-field cron expression and an IANA time zone — 0 8 * * * with America/New_York means 8:00 every morning in New York, including daylight-saving changes.
Schedules live on the server and keep firing with the browser closed. Pause one without deleting its configuration; when re-enabled, the next run time is calculated from the same expression and time zone.
Orchestration
An orchestration chains complete pipelines. It uses the same canvas-and-chat layout as a data pipeline, but every card represents a whole pipeline. Connect an output handle to the next pipeline's input to define execution order; the pipeline at the head of the chain carries the schedule.
Each connected pipeline waits for its predecessor to succeed. If a step fails, the chain stops before incomplete data reaches downstream work.
Storage
The Storage tab creates a durable managed relation immediately — there is no separate infrastructure or migration step. Resources belong to the open workspace and can be used by its Database Connector blocks.
- Table
- An
id, a JSONBdatacolumn, and timestamps, for structured pipeline records. - Object store
- An object key, binary content, content type, JSONB metadata, and timestamps, for files or raw payloads.
Each resource has a table view and a read-only SQL editor scoped to that resource, where it is addressed as this_table. Deleting a storage resource asks for confirmation, because it drops the underlying relation and its data.
MCP servers
A workspace MCP server exposes managed storage to your own agents through an authenticated Streamable HTTP endpoint. Connect workspace tables, add focused server instructions or prompts, create an access token, and use the generated URL from an MCP client.
The chat can connect named internal resources, connect all workspace storage, remove nodes, or pause the server. Hosted MCP connections are always read-only Streamable HTTP.
Credentials stay separate. External nodes hold connection metadata only. Put passwords, tokens, and API keys in the Key Vault — never in endpoints, chat messages, or node names.
Key Vault
The Key Vault stores external-service credentials encrypted at rest. Values are masked by default and revealed only deliberately. Scope decides which pipelines may use a key:
- Workspace key
- Reusable across pipelines, linked or unlinked from individual pipeline rows.
- Pipeline only
- Created for and isolated to a single pipeline.
Web collection and managed storage use Pipechat's own service credentials. Every block that calls an LLM provider requires a compatible API key from your Key Vault — Pipechat never supplies or bundles model usage.
Use the narrowest scope that fits, rotate credentials at the provider, then update the stored value. Never paste secrets into chat messages, prompts, code files, or run logs.
Desktop
The desktop app includes the same interface as the browser app, while collection, enrichment, pipeline Python, and data storage run on your own machine instead of Vercel Sandbox. Pipeline data stays in a PostgreSQL database on your computer.
Python 3 and PostgreSQL are required and checked on first launch. Internet access is only needed for the external websites and model providers a pipeline uses. AI blocks stay locked until you connect your own LLM API key in the encrypted Key Vault.
Signed-in users can download the installer for their current platform from Account Settings.
