AI Builder
Pipe, our open-source fork of the Pi harness, turns a conversation into a working data pipeline you can inspect, edit, and run.
Describe the site you need. Pipechat builds and validates reliable extraction rules against real captured pages, then runs them deterministically to structure, enrich, and deliver consistent web data to your database or agents.
Pipe, our open-source fork of the Pi harness, turns a conversation into a working data pipeline you can inspect, edit, and run.
Pick from a growing list of maintained sources and a library of premade pipeline recipes — run them as they are, or tweak them to fit your setup.
Connect the models that fit the job — from paid LLM providers to Hugging Face, Groq, OpenRouter, and self-hosted open-source models.
LLMs monitor and repair external connections when sites change — whether you use a Pipechat-maintained source or connect your own source from the open internet.
Use hassle-free managed cloud compute for scheduled runs, or run the same pipeline locally on your own machine and infrastructure.
Host your data behind an MCP server and query it directly from AI chats and agents such as ChatGPT and Claude.
SOURCES AND RECIPES
The Source Library holds sources we keep current for you — address, extraction rules, validation evidence, and the full version history. Connect one to a pipeline and it behaves like any other block, except you never wrote it and you never have to fix it. Recipes go a step further: prebuilt pipelines wired end to end that you can run as they are, or tweak into your own.
MODEL AND STORAGE AGNOSTIC
Pipechat is the layer between the two, not a replacement for either. The model that reads a page and the store that keeps the records are both settings — change one and the same pipeline keeps running, on your keys and inside your own infrastructure.
LOCAL OR CLOUD
Use Pipechat Cloud when you want managed compute and scheduled runs, execute locally inside the desktop app, or download the complete codebase and operate it yourself outside Pipechat.
# The same project runs in Cloud, Desktop, or your own environment. from __future__ import annotations import os from pipeline_runtime.context import as_records from pipeline_runtime.models import complete # Provider keys and runtime target stay in environment config. API_KEY = os.environ["ANTHROPIC_API_KEY"] RUNTIME = os.getenv("PIPECHAT_RUNTIME", "local") def run(payload, context): records = as_records(payload) for record in records: record["industry"] = complete( api_key=API_KEY, model="claude-sonnet-5", prompt=f"Classify this company: {record['body'][:2000]}", ) return records
Open a workspace to see the pipelines inside it.
this_table as its name.No prompts yet. Add one for a repeatable way to use these datasets.
Run #00000000
Custom Code
MANAGED SOURCE BUILDER
Choose a generated file to inspect it here.
No snapshot selected.
Reusable sources your workspace can connect to from any pipeline.
External Sources / Connections
Primary company website used to collect company, product, and positioning information.
Pages and refresh rules for this connection
Completed 4 minutes ago
Content is current and available to downstream enrichment blocks.
Fields this source provides to the published dataset
Pipechat managed source
Pipechat monitors and updates this source. You can review its version and health history without exposing the underlying extraction code.
Source extraction code
Versions / V3
Extraction monitoring
Health Check / Result
25 records were pulled and validated against the dataset schema.
Preview of records produced by V3
| company_name | domain | industry | employee_range | source_url |
|---|
Select a file to view its source code.
Run history / #1042
All four stages completed and the dataset was published.
Execution details for each pipeline block
Important events from this execution
Hosted runtime
Configure collection limits and link your own model-provider API key to every AI block.
The Run button uploads the generated Python project to a fresh Sandbox and saves input/output snapshots for every block.
Pipeline automation
Use standard five-field cron syntax. Times are interpreted in the selected time zone.
minute hour day month weekdayTriggers are stored on the server and run even when this page is closed.
Add a cron trigger above to automate this pipeline.
Template pipelines you can inspect and adapt to your own sources, models, and destinations.
Encrypted by Supabase Vault for the whole workspace or one pipeline.
A detailed guide to building, running, automating, and managing data in Pipechat.
Pipechat is a visual data-pipeline workspace. You describe the dataset you want, assemble the work as connected blocks, and run the generated project on managed compute now or on a schedule. The graph, its code, its data, and its execution history stay connected instead of living in separate tools.
The usual path is create a workspace → add a pipeline → connect and configure blocks → run and inspect → automate. You can do the mechanical work directly in the builder or ask the chat agent to change the graph and its code.
Pipechat persists blocks, connections, canvas positions, generated code, chat messages, named versions, run snapshots, schedules, orchestration state, managed storage resources, and encrypted keys. Reloading the page or closing the browser does not discard the project.
Supabase manages durability. The application database, workspace tables, and object records live in the connected Supabase PostgreSQL project. A destructive action such as deleting a storage resource still asks for confirmation because it drops the underlying relation and its data.
The Workspaces page is the front door. Its rows summarize how many pipelines, storage resources, orchestrations, and external sources belong to each workspace. Open a row to keep the workspace name in the breadcrumb while switching among three matching catalog views:
Use + New Workspace at the root and the context-sensitive + New… action inside a workspace. A workspace can only be deleted after its pipelines are removed, which prevents accidental orphaning.
The Pipeline tab is the visual source of truth. Blocks run in connection order: records leave a source, move through transforms, and arrive at a destination. Drag blocks to make the graph readable; connect an output handle to the next block's input; select a block to inspect its configuration.
The top bar keeps the workspace, pipeline name, current version, and run action together. Changes are persisted as you work, while versions provide explicit restore points.
A connection is both visual and executable: it determines which upstream output becomes the next block's input. Keep a single clear direction through the graph unless a block deliberately fans records into multiple branches.
The chat panel can add, update, connect, move, and remove blocks. Describe the result and constraints — for example, “collect press releases, extract company and funding amount, then store unique records” — and the agent updates both the graph and its isolated project files.
Code Files exposes the generated project behind the canvas. Each block maps to readable files, so you can inspect how it works, edit precise behavior, and verify the agent's changes. Graph edits and file edits remain part of the same pipeline project.
Be specific about input shape, required output fields, deduplication rules, and failure behavior. Those details give the agent a much better contract than a broad request such as “clean the data.”
The version menu beside the pipeline name shows the active snapshot. Quick save records the current state; a named save creates a recognizable checkpoint before a larger change. Use the recent-version list to restore an earlier graph and its code.
Run pipeline materializes the current project, starts a managed execution, and records the outcome. Run History combines manual, scheduled, and orchestration-triggered executions in one place.
Every row shows status, trigger, start time, duration, and output summary. Open it to move through individual blocks and compare what entered a step with what came out.
Start at the first failed block, not the final destination. Read its error and log, inspect its input snapshot, then compare that shape with what its code expects. Common causes are a changed source page, a missing field, an unlinked key, an invalid response from an external service, or a compute timeout.
After making a fix, run the pipeline again. Pipechat preserves the failed attempt and creates a new run, giving you a clean before-and-after trail instead of overwriting the evidence.
The Scheduling tab creates recurring triggers for one pipeline. Enter a standard five-field cron expression and an IANA time zone. For example, 0 8 * * * with America/New_York means 8:00 every morning in New York, including daylight-saving changes.
Schedules live on the server and continue to fire with the browser closed. Pause one without deleting its configuration; when enabled again, the next-run time is calculated from the same expression and time zone.
Choose Create Pipeline, switch the modal to Orchestration, and name it. Pipechat opens a full builder that uses the same canvas-and-chat layout as a data pipeline, but every card represents a complete pipeline.
Open Add pipeline to browse the workspace pipelines in a table and place the steps you need, or describe the flow in chat. New cards begin unconnected. Connect an output handle to the next pipeline’s input to define execution order, and drag cards only to arrange the canvas. The orange pipeline at the head of the connected chain is the scheduled start.
Click the orange starting pipeline to open its compact schedule modal, then set a five-field cron expression, IANA time zone, and enabled state. Each connected pipeline waits for its predecessor to succeed; if a step fails, the chain stops before incomplete data reaches downstream work.
The chat can add or remove named pipelines, put a pipeline first, build the whole workspace in order, and translate requests such as “weekdays at 8 AM in America/New_York” into the starting schedule. Review the graph and press Save orchestration when the flow is ready.
The orange type tag distinguishes an orchestration from a blue Data Pipeline row. Its step count, last run, updated time, and enabled or paused status use the same catalog columns. Open the row to return to the builder; turn off Enabled to pause it without losing the chain.
The Storage tab creates a durable managed relation immediately; there is no separate infrastructure or migration step. Resources belong to the open workspace and can be used by its database connector blocks.
id, JSONB data, and timestamps for structured pipeline records.Select a resource row to browse up to 100 recent rows. The SQL editor is read-only and scoped to the open relation; refer to it as this_table. Renaming keeps the resource's identity in Pipechat. Deleting asks for confirmation, then drops the physical relation and all of its contents.
The MCP tab is a catalog of servers in the open workspace. Create as many servers as the workspace needs; each row keeps its own endpoint, transport, enabled state, and storage graph.
Open a row to enter its builder. One MCP server block sits in the center holding server instructions and prompts, and connects out to storage tables from this workspace. Drag any block to arrange the graph; positions are saved with the server.
The hosted endpoint uses Streamable HTTP and exposes read-only discovery, structured query, search, and fetch tools. The Settings tab holds server instructions, named prompts, the generated URL, and access-token controls. The chat can connect tables, remove blocks, change their configured access, or pause the server. Edits save on their own once they settle.
Credentials stay separate: an MCP server exposes tables Pipechat already manages. Put passwords, tokens, and API keys in the Key Vault rather than endpoints, chat messages, or block names.
The Source Library contains reusable, maintained sources. A source records the address, publisher, availability, extraction rules, and generated files needed to collect listings and individual pages. Review a source before linking it so its output shape matches the pipeline that will consume it.
Availability tells you whether the source is ready to collect. Connection state tells you whether it is already attached to the current pipeline. The library definition is reusable; each pipeline run still fetches current content.
Supabase Vault stores external-service credentials with authenticated encryption at rest. Pipechat stores only Vault references and display metadata. Values are masked by default and are only revealed deliberately. Scope and links decide which pipelines are allowed to use a key:
Web collection and managed storage use Pipechat’s service credentials. Every block that calls an LLM provider requires a compatible API key from your Key Vault; Pipechat never supplies or bundles model usage.
Safe operating habit: use the narrowest scope that fits, rotate credentials at the provider, then update the stored value. Never paste secrets into chat messages, prompts, code files, or run logs.
Account