Skip to content

Repository files navigation

polyflow

A workflow engine for AI agents. The agent runs a small state machine instead of deciding one tool call at a time. polyflow loads that workflow only if every path through it satisfies rules you wrote, then hands the agent one instruction at a time and keeps the run's state outside the conversation.

It ships as an MCP server, so any MCP-capable agent — OpenWorker, Claude Code, Kiro, Hermes, DeepSeek Harness — can use it with no changes to that agent's core.

Experimental, unproven, not peer-reviewed. The check is a consistency check, not a proof, and "exhaustive" always means exhaustive over the finite domain the contract declares. Every finding is a lead, not a result.

polyflow is the single-participant engine: one process, one broker in memory, stdio, SQLite, nothing to install. For several sessions or several people working one run at the same time — a shared broker, orders they claim, and a page a person can watch — see polycrew, which uses polyflow as a library and adds two tools to these six.

What it is for

Work that repeats, outlives one session, and ends in something you cannot take back. A scheduled job that posts or files or charges. A deploy. Anything where doing it twice is a real problem and where you can write down a rule in one sentence.

Two results from testing it against two different agent harnesses (FINDINGS-phase3.md, 48 runs, records in runs/):

  • A scheduled job that fires twice in one day — what a scheduler does after downtime — repeated the whole job and posted twice in 8 of 8 runs without polyflow. With it, 0 of 8: the second session asked for the workflow by name, saw it had finished, and stopped. Same split on both harnesses.
  • No run in any condition posted without an approval, 48 out of 48 — with polyflow and without it. On a single clean run the engine changes nothing about safety. What separates them is the second run.

The reason it holds: the workflow is not something the agent remembers. It is not compacted, summarized, or retrieved into a prompt. It sits in a database and the agent queries it. The second session did not recall that the job was done — it asked.

The loop

tools → observe → reason → WORKFLOW ──▶ polyflow admits it (or refuses)
                                            │
                        ┌───────────────────┘
                        ▼
        one work order  →  the agent runs the tool, through its own
                           permission gates, with its own credentials
                        →  workflow_report
                        →  next work order … until terminal

The agent never decides what comes next. It reasons about how to fulfil one order — which is what a model is actually good at — and reports the result. Sequencing, retries, timers, duplicate suppression and terminal conditions belong to the machine.

Rules are checked before a workflow can run

A workflow ships with invariants: rules that must hold on every path it can take, written as small predicates.

{ name: 'no-post-without-prior-approval',
  pred: (path) => path.emitted.every((e, i) =>
    e.kind !== 'post_brief' || path.actionBefore('APPROVED', i)) }

Startup enumerates every reachable emission path over the contract's declared domain and checks them. A workflow that fails is not registered — not flagged, unrunnable:

[polyflow] admitted: customer-brief — paths explored: 5 · states seen: 10 · exhaustive within declared domains
[polyflow] REFUSED: unsafe-brief
[polyflow]   no-post-without-prior-approval

test/fixtures/unsafe-brief is the deliberately broken twin: it posts on entering review, before the human answers. It still calls ask_user, still targets the same channel, still satisfies a standing grant on that tool. A reviewer reading the diff could easily miss it. The gate does not.

Writing a checked rule beats a standing grant for the same reason. An unattended automation is normally approved by verb — "allow slack_send to #cs", forever, for whatever the model decides to do with it. That is the ceiling when the plan is a prose instruction re-planned on every run.

Quickstart

npm install                    # pulls polygraph (polyrun) as a dependency
npm test                       # 19 tests, no API key, deterministic
node bin/polyflow-mcp.mjs      # MCP stdio server

Running alongside OpenWorker

Prerequisites: Node 22.13 or newer — polyflow uses node:sqlite, which needs --experimental-sqlite on earlier 22.x — and OpenWorker installed. The engines field enforces the floor at install time. polyflow needs no API key of its own: it never calls a model.

1. Register it. From the polyflow directory:

node bin/polyflow-install.mjs --agent openworker/cowork --workspace acme
# --print shows the entry and the target path without writing anything

This merges a polyflow entry into OpenWorker's global mcpServers file — the same one the Connectors page edits (%APPDATA%\coworker\mcp.json on Windows, ~/.config/coworker/mcp.json otherwise, $COWORKER_STATE_DIR overriding both). It merges rather than replaces, and refuses to touch a file it cannot parse.

2. Restart OpenWorker. There is no polyflow daemon to start or supervise: OpenWorker spawns bin/polyflow-mcp.mjs over stdio when a session opens and tears it down with the session. Run state lives in the SQLite file at POLYFLOW_DB, so it survives both.

3. Check it came up. The six tools appear as mcp__polyflow__*. Ask the agent to "list the workflows you can run" — it should come back with customer-brief, its admitted: true, and the five guarantees it was admitted under. If it does not, the Connectors page carries the standing error, and the server's own startup lines (admitted: / REFUSED:) go to stderr.

4. Use it. Nothing special: give the agent a task a workflow covers and it picks the workflow up on its own — that is what FINDINGS-phase3.md measures. To put a recurring job on it, create an ordinary OpenWorker automation whose instructions describe the task; the workflow re-attaches by derived key on every fire instead of starting over.

Areas. --agent is the agent-class area (which library of workflows this kind of agent draws on) and --workspace is the instance area (whose runs these are).

One install can serve several workspaces, but each needs its own entry under its own key — a second run of the installer with the same --name replaces the first:

node bin/polyflow-install.mjs --workspace acme
node bin/polyflow-install.mjs --name polyflow-beta --workspace beta --db ~/.polyflow/beta.sqlite

Point both at the same --db to share a store, or at different files to keep the runs apart.

Where state lives. Run from a clone, the store is .polyflow/polyflow.sqlite and the library is ./workflows. Installed as a package, both default to ~/.polyflow/ instead — a global install lives under node_modules, which the next upgrade deletes and re-extracts, and durable runs must not be in there. --db and --workflows override either.

Adding your own workflow. Copy workflows/customer-brief/ into your library directory (POLYFLOW_WORKFLOWS, printed by the installer) and edit the six files (see A workflow below). Restart the server: a workflow that fails its emission check is refused at startup and cannot be started at all, so a bad edit fails loudly rather than at 3AM.

Permissions. The installed entry sets requires_approval: false deliberately — polyflow tools reach nothing outside the machine, and the run's real side effects are the agent's OWN tools, which keep their own gates. Prompting on every workflow_report would put a dialog between the agent and its own bookkeeping. The entry also declares tool_risk for the read-only tools, honoured with upstream/0001-mcp-per-tool-risk-level.patch applied and harmlessly ignored without it.

Where polyrun comes from. polyflow embeds polyrun in process, resolved from node_modules/polygraph, then a sibling checkout. Setting POLYFLOW_POLYRUN overrides both, and a value that does not contain a polyrun build is an error rather than a silent fall-back to the packaged one.

Other agent hosts

polyflow is a plain MCP stdio server, so anything that speaks MCP can use it. The installer writes the right file for each host:

node bin/polyflow-install.mjs --host kiro          # ~/.kiro/settings/mcp.json
node bin/polyflow-install.mjs --host kiro --scope workspace   # ./.kiro/settings/mcp.json
node bin/polyflow-install.mjs --host hermes        # ~/.hermes/config.yaml
node bin/polyflow-install.mjs --host claude-code   # ./.mcp.json
node bin/polyflow-install.mjs --host generic       # prints the entry, writes nothing

Two hosts take a different shape and are printed rather than written:

node bin/polyflow-install.mjs --host dsh       # cordis.yml plugin entry for DeepSeek Harness
node bin/polyflow-install.mjs --host nemo      # YAML for a NeMo Agent Toolkit workflow
node bin/polyflow-install.mjs --host registry  # AWS CLI call to publish an Agent Registry record
  • Kiro / Kiro Crew reads mcpServers from ~/.kiro/settings/mcp.json (user) or .kiro/settings/mcp.json (workspace, which wins on a name clash). Kiro Crew's recurring unattended jobs are the same shape as OpenWorker's scheduled jobs, which is the case the results in FINDINGS-phase3.md are about.
  • Hermes Agent keeps its MCP servers in ~/.hermes/config.yaml under mcp_servers:, alongside the rest of its configuration. There is no YAML parser here and none is wanted — round-tripping a live config would lose comments and formatting — so the edit is textual and narrow (src/yaml-block.mjs, unit-tested in test/yaml.test.mjs): it replaces the polyflow entry and nothing else, keeps the previous file as .bak, and writes by rename. HERMES_HOME overrides the location. The entry sets trust: untrusted, which is the right default for a server you did not write: Hermes then prompts on every call that is not annotated read-only, and polyflow's three read-only tools are annotated, so browsing a run costs nothing while every write still stops for approval. polyflow is stdio, so there is no hermes mcp login step. If you already have polyflow in ~/.claude.json, hermes import-agent claude-code migrates it.
  • DeepSeek Harness (dsh) bridges MCP servers through its dsh-mcp-client plugin, one instance per server, configured in cordis.yml. That file is a list of plugin entries rather than a map keyed by server name, so there is nothing to merge into safely without a YAML parser — the entry is printed and you append it. It is set to failOnStartupError: true, so a broken polyflow fails activation instead of leaving the agent with no workflow tools and no sign of why. dsh validates structuredContent against a tool's advertised outputSchema, which polyflow declares for all six tools.
  • NVIDIA NeMo Agent Toolkit connects through its mcp_client function group (needs nvidia-nat-mcp). The printed block declares the group and adds it to a workflow's tool_names. NeMo can also run as an MCP server itself, so a NeMo workflow can be one of the tools a polyflow work order names.
  • AWS Agent Registry is a catalog rather than a runtime: publishing a record lets other people and agents in the organization discover polyflow. Records can be synchronized from an HTTPS endpoint, which a stdio server has no way to offer, so the printed command creates a manual MCP record instead.

The registry and polyflow both gate on approval, but not the same approval. A curator approves that a server may be found; polyflow's admission check decides that a workflow may run. Those are different questions, and an organization can use both: publish polyflow in the registry so teams can discover it, and let polyflow refuse the workflows that break their own rules.

Only the OpenWorker path has been exercised end to end (see FINDINGS-phase2.md). The others are built from each host's documented configuration format and have not been run.

Env vars, whichever host you use:

env meaning default
POLYFLOW_WORKFLOWS workflow library directory ./workflows
POLYFLOW_DB sqlite path .polyflow/polyflow.sqlite
POLYFLOW_AGENT agent-class area default
POLYFLOW_INSTANCE instance area (workspace) cwd basename
POLYFLOW_POLYRUN polygraph checkout ../polygraph

Tools

tool does
workflow_list what this agent knows how to do, and the guarantees each was admitted under
workflow_start start or re-attach — the run's identity is derived from validated input, so a nightly task resumes instead of restarting and an agent cannot rename its way to a second run
workflow_report report a tool result, receive the next order
workflow_state state + open orders, changes nothing
workflow_signal an out-of-band event; an action that does not apply is an observable reject
workflow_journal every step, accepted or rejected, with its reason — also a valid Polygraph trace corpus

Areas

Two tiers, and they need no new fields in OpenWorker:

  • agent area — one per agent class (openworker/cowork). Owns the workflow library: what this kind of agent knows how to do. Maps to ScheduledTask.agent.
  • instance area — one per running copy (workspace). Owns the live runs and their journals. Maps to workspace, which is already coworker.memory.Scope.WORKSPACE.

The instance id is derived from agent | instance | workflow | key, which is why start and attach are one call.

A workflow

Six files in a directory:

polyflow.workflow.json   name, area, tools{effect kind -> agent tool},
                         key{template,fields} — the run's identity, derived
contract.json            states, actions, finite data domain
machine.cjs              SAM v2 strict-profile module
effects.cjs              pure mapper: transition -> work orders
effects.manifest.json    completion actions + retry policy per kind
effect-invariants.mjs    what may be EMITTED, on every reachable path

The inversion that makes this work for agents: in polyrun the runtime executes effects. polyflow has no credentials, no connectors and no permission engine — the agent has all three. So an effect is a work order handed back. The handler parks; the agent claims the order, runs the tool under its own gates, and reports. Only then does the completion action dispatch.

Durability falls out of the lease machinery. The pending map is in-memory, so a crash loses the promise, the lease expires, the effect is re-claimed and the order is re-offered — same intent id, at-least-once, absorbed by the machine.

What the tests prove

✔ the admission gate certifies the demo workflow exhaustively
✔ workflow_list reports the guarantees the run was admitted under
✔ happy path: one order at a time, ending posted
✔ the run key is derived from input, not chosen by the caller
✔ an invalid key field is refused with an instruction, not honoured
✔ a finished run says so, and says not to start another
✔ start is idempotent: re-attaching returns the run in progress
✔ a denial is a result, not a fault — and no post is ever ordered
✔ zero tickets ends the run rather than posting an empty brief
✔ a duplicate report is refused, not double-executed
✔ an out-of-band action that does not apply is an observable reject
✔ a workflow that can post before approval is REFUSED and cannot be started
✔ a run outlives the process: restart re-offers the open work order
✔ initialize, tools/list, tools/call over stdio
✔ an empty file becomes a config with just our block
✔ a config with no mcp_servers keeps every line it had
✔ other servers survive, comments and all
✔ re-running replaces our entry and only ours
✔ values are quoted, so a Windows path is not read as YAML syntax

The restart test is the one that matters: session 1 drives the run to the approval step and dies; session 2 is a different process with no conversation, no transcript and no replay — because the state was never in the messages to begin with. It picks the run up exactly where it was, and exactly one post happens across both.

Not built yet

  • Promotion. Workflows are hand-authored here. The plan is to mine recurring run shapes out of journals and propose a machine for review — induction from history, not foresight. Authoring a machine per task costs more than the tool calls it replaces unless it is reused.
  • Versioning. polyvers gates a changed workflow against in-flight runs; not wired in.
  • Audit. The journal is already a trace corpus; polyrun audit against it is not wired in.
  • Signed journals. Hash-chain the journal rows and sign the digests, the way Dapr 1.18 signs workflow history, so a run's record can be verified by someone who did not produce it. Cheaper here than in a replay engine: the journal is a record, not the resumption mechanism, so signing it imposes none of the determinism discipline replay forces. Worth pairing with the audit above — a signature proves the record was not altered, an audit proves it is consistent with the machine that claims to have produced it, and they are different claims.
  • The OpenWorker seams that need core changes: routing a parked order to the Inbox, and workflow_ref on ScheduledTask. See FINDINGS-phase0.md.

About

A workflow engine for AI agents: the agent reasons a workflow, polyflow admits it only if it model-checks, then runs it durably and hands the agent one work order at a time. Ships as an MCP sidecar.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages