PtcRunner

A runtime for building self-improving agents

Build loops that check their work and recover from evidence.

Bound execution, inspect failures, verify outcomes, and test changes before adopting them.

01
A language for agents

Small, typed, and limited to the tools you approved, discovered in a REPL rather than listed in the prompt.

02
A runtime for agent services

Built for bounded, traceable jobs and many concurrent requests, without a desktop, workspace, or container for every agent.

Three runnable examples

See the runtime solve real workflow problems.

One demo pattern

Generated code is tested before it becomes the parser.

This example uses one workflow and three mission environments to show code evaluation and permissions. Each environment gets different code, data, and tools. It is one of many patterns you can build with the same runtime.

First run repair and verify

  1. Artifact missionLook for accepted.cljIt is missing.
  2. Browser missionRun the installed parserOld selectors return no records.
  3. Evidence missionAsk for a selector recipeToy evidence functions expose the failure and one changed page.
  4. Artifact missionWrite candidate.cljThe candidate is real PTC-Lisp, not trusted host code.
  5. Browser missionEvaluate it twiceRun on the changed page and a page the model did not see.
  6. WorkflowAdopt only if both passThe verified source becomes accepted.clj.

Outcomes drive the workflow. Each evaluation returns a value or a bounded failure. The workflow can inspect that outcome, evaluate another program, retry, compare results, or adopt a candidate. The full trace is recorded separately; reading it inside a workflow requires an explicit capability.

Next runaccepted.clj → browser mission → result No model call.

Step through the full runtime flow →

Why not a coding agent?

Great tools, built for a different job.

An agent's results come from the model and the harness around it. Coding agents have the strongest harnesses today, so people use them for many non-coding jobs too.

But a service has different needs. It must handle many requests, stop one run from using too much time or memory, and explain later why it made a choice.

PtcRunner makes those runtime features: fixed limits on time, memory, and tool calls; a structured trace from every run; and thousands of isolated runs on one machine. Run workflows with ptc run, then inspect their traces from another authorized PTC-Lisp program, the REPL, or the Viewer.

Change the agent, not its authority

The loop is a library, not the runtime.

A system that improves itself must be able to change itself. In many frameworks, the agent loop is trusted host code. Changing the loop can also change what the agent may do.

Here the loop is an ordinary PTC-Lisp library. You can replace how the agent works without giving it more tools. Replay holds the model fixed, so you can test a candidate against the old version before a person decides whether it ships.

The loop is a library, not the runtime. A workflow prelude, compiled from selected
            components, contains a stack of agent.main, agent.core, agent.prompt, agent.retry and
            llm, outlined and tagged 'replaceable'. In the centre a four-step ring driven by
            agent.core: prompt, model writes PTC-Lisp, evaluate, observation. On the right a
            separate mission prelude holding the generated program, your domain components, and
            prompt-visible exports. A bar across the bottom reads: every tool: requirement is
            checked against the assembled providers — never granted by them.
The workflow and mission are separate. A library may ask for a tool, but only the host can grant it.

Runtime guarantees

Small programs. Clear limits. A record of every run.

Bounded by design

No import, open, fetch, or shell. A sandbox takes a language that can do anything and removes the dangerous parts one by one. Here there is nothing to remove: a program can use its input and the tools you approved, and nothing else.

Checked at the boundary

Components and tools carry signatures. Inputs are validated before a call runs and outputs after it returns, so a value of the wrong shape stops at the boundary instead of flowing on.

Trace every run

Inputs, tool calls, evaluations, outcomes, and limits become structured evidence you can inspect and test.

ptc-host.json installs providers and sets outer limits. ptc.json selects from them and may make the limits tighter for one project.

Will the model write it?

Yes. Give it a small language and a narrow job.

PTC-Lisp is a bounded, Clojure-like language with types checked on input and output. The model writes one short program per turn instead of a long chain of tool calls, the pattern often called code mode, against tool signatures it can explore in a REPL.

The tutorials use a small, low-cost model by default. Try it on your tasks and measure the result.

Take the language tour →

Improve from evidence

Check an answer now. Test a better workflow for later.

Evidence is used at two different times. During a run, the DABStep example makes two analyses and a reviewer agree before the workflow returns an answer. After a run, the repair example reads the failure's trace, proposes a change, and tests it on inputs the model did not see.

Code travels; authority stays. A wide box labelled WORKFLOW ENVIRONMENT,
            host-authored code, the only sequencer, holding preloaded.clj (before turn 1)
            and the agent loop (the model writes PTC-Lisp). Navy arrows labelled
            kernel/eval-with descend into two mission boxes, debug.nav with keys to read
            the failed run's traces and failed-run-traces with keys to raw trace pages,
            and dashed arrows return the evidence packet, composed into the turn-1 prompt.
            An orange arrow labelled model-written code descends into a third mission box,
            synthesize, whose keys are repair.terminal only, and a return arrow carries a
            typed decision: propose or abstain. A legend maps navy to host-authored code
            and orange to model-generated code: same evaluator, same trace. A bar across
            the bottom reads: evidence flows down into the prompt — authority never flows
            with it.
Evidence can move into the model's prompt. The tool authority stays with the workflow that gathered it.

A repair run is an ordinary run too. It leaves a trace, so the debugger can be tested and improved with the same tools it uses on everything else.

Optional local viewer

Inspect a run when you need to.

The runtime records structured traces without needing a UI. The included Viewer is a convenient way to explore one: see the result, limits, tool calls, model sessions, and errors in one place.

Open the Viewer reference →
PtcRunner Viewer showing a successful deterministic run of the orders tutorial:
              the run summary with its duration and bundle, the JSON result, and counters for
              evaluations, capability calls, model calls, errors, and events.
A deterministic tutorial run, shown in the current Viewer.

Built on the BEAM

Many isolated runs. One small runtime.

The BEAM is the virtual machine built three decades ago for telephone switches and hardened in production ever since, where huge numbers of concurrent connections, low latency, and failure isolation were the requirements.

Each environment runs in its own BEAM process and gets what an operating system would give a program: preemptive scheduling, so a runaway job cannot starve its neighbours; a per-process heap limit enforced by the VM, so a run over its memory budget is stopped without taking the machine; and let-it-crash isolation, so one failing run never corrupts another. Processes are cheap enough that one machine hosts thousands of runs, with no container per run.

Try it locally

One executable. No separate runtime.

Download the self-contained macOS arm64 archive from GitHub Releases, or pull the Linux Docker image.

0.x, under active development. Breaking changes are expected.

terminal
ptc init hello-ptc
ptc run hello-ptc/ptc-project.json

{"greeting":"hello world"}

Runs offline and writes a structured trace.