Caravan

Caravan is a typed agentic CLI harness and LLM orchestration framework for OCaml — a working autonomous agent out of the box, light enough for HPC nodes, containers, and non-root environments.

❯ /agent summarize the files in this directory
⏺ ls({"path": "."})
  ⎿ # ls .  (cwd: /home/you/project) (+14 lines) [0.0s]
  ✔ Task finished: An OCaml project with a dune build, 3 libraries…

Why Caravan

AccountableEvery model call, tool call, and nudge is a structured event; sessions write JSONL transcripts to ~/.caravan/logs/.
13 providersFrom a 1B llama on a laptop to Claude / GPT-4o / Gemini, behind one interface.
GovernedTool permission modes auto / ask / readonly; verified TLS.
Scripting-nativecaravan agent "task" --json → one JSON object, real exit codes.
HygienicOne binary, one 0600 TOML config, no runtime file spew.
TypedPipelines are 'a -> ('b, string) result; tools are typed first-class modules.

Where to go

Installation

Requirements

  • opam ≥ 2.1 — the only hard prerequisite
  • OCaml ≥ 5.1 (the installer creates a 5.2 switch if needed)
  • system libraries: libssl-dev, libgmp-dev, pkg-config

Everything installs into ~/.opam — no root required.

One-liner

curl -fsSL https://raw.githubusercontent.com/adukhan99/Caravan/main/scripts/install.sh | bash

The installer is idempotent: it initializes opam if needed, ensures an OCaml 5 switch, clones (or updates) the source in ~/.caravan/src, builds, and installs the caravan binary onto opam's PATH. Override the clone location with CARAVAN_SRC, the repo with CARAVAN_REPO.

Manual build

git clone https://github.com/adukhan99/Caravan.git && cd Caravan
opam install . --deps-only --with-test -y
dune build && dune test
dune exec caravan -- init

First run

caravan init      # pick a provider, model, API key (input hidden; config saved 0600)
caravan doctor    # verify config, keys, endpoint reachability
caravan           # chat

Upgrade / uninstall

# upgrade: re-run the installer
# uninstall:
rm -rf ~/.caravan
opam remove Caravan

Quick Start

Chat

caravan

Type to chat. / opens the live command palette — Tab completes. Arrows edit; Up/Down recall history.

Let it work autonomously

Inside the REPL:

/agent find the failing test in this repo and explain why it fails

Or from a script / batch job:

caravan agent "profile the hot loop in sim.c and propose a fix" \
    --max-turns 20 --json | jq -r .result

Exit code 0 on success, 1 on failure or exhausted turn budget. --json returns the result, token usage, and the transcript path.

Switch models

caravan -p ollama -m llama3.2:1b        # tiny local
caravan -p anthropic                    # frontier (uses provider default model)
caravan providers --ladder              # a curated model per weight class

The browser cockpit

caravan web        # http://127.0.0.1:8787 — localhost only

Chat, run agent tasks with a visible tool audit trail, and edit every setting (including API keys) from the ⚙ settings panel.

Guard rails

caravan config set permissions ask      # confirm before mutating tools
CARAVAN_PERMISSIONS=readonly caravan agent "audit this repo"   # look, don't touch

CLI Reference

Every command accepts -p/--provider, -m/--model, --base-url, -s/--system. Resolution order: CLI flag → CARAVAN_* env → ~/.caravan/config.toml → registry default.

Commands

CommandPurpose
caravaninteractive REPL (default)
caravan agent "<task>"one-shot autonomous run (run is an alias)
caravan complete "<prompt>"single completion, no tools
caravan web [--port N]local web UI (127.0.0.1 only)
caravan providers [--ladder]provider table / model ladder
caravan modelsmodels on the current provider
caravan config show|path|keys|get K|set K Vinspect or edit the config
caravan mcp add|remove|list|getmanage Model Context Protocol (MCP) servers
caravan initsetup wizard
caravan doctor [--fix] [--json]diagnostics, with fixes offered (exit 1 on failure)

caravan mcp

  • caravan mcp list — list configured MCP servers and their health status
  • caravan mcp get <name> — inspect an MCP server and its discovered tools
  • caravan mcp add <name> [--transport stdio] [--no-probe] -- <command> [args...] — probe connection, register tools, and save to config
  • caravan mcp remove <name> (or rm) — remove an MCP server configuration

caravan agent

  • --max-turns N — turn budget (default: config max_turns, else 24)
  • --quiet — final result only
  • --json — one JSON object on stdout: {ok, result, turns, usage, transcript}
  • exit codes: 0 done · 1 failed / out of turns · 2 config error

REPL slash commands

Typing / opens a live palette; Tab completes the command, and once one is typed, its arguments — /config set <Tab> offers the settings, /provider <Tab> the registry. Highlights:

CommandEffect
/agent <task>autonomous loop with tools
/resume [path]restore session state from a saved checkpoint
/nudge <text>steering note injected before the next model call
/lisp <program>evaluate a Slip expression
/model · /models · /provider · /providersswitching
/subagents · /subagents add · /subagents remove [name]worker roster and CRUD
/permissions [mode]auto | ask | readonly, live
/config · /config keys · /config set k v · /config unset k · /config get k · /config editsettings
/key <provider>store an API key (hidden input)
/system /temp /top_p /top_k /max_tokens /seed /stopgeneration
/memory <n> · /summarisecontext window / compact now
/history · /export [file] · /toolsinspect the session
/mcp [list|add [--no-probe]|get|remove]manage MCP tool servers and dynamic tool bindings
/plugins [enable|disable <id>]plugin composition and lifecycle states
/doctor · /initpre-run commands, callable in-session
/clear · /help · /quithousekeeping

Line editor

KeysEffect
alt+enter · ctrl+oinsert a line break — enter submits
ctrl+rsearch history; enter runs the match, esc cancels
tab · ↑ ↓complete a command or its argument
↑ ↓move between rows, or walk history
ctrl+a ctrl+e · ctrl+k ctrl+u ctrl+wline start/end · kill to end, to start, previous word
ctrl+c · ctrl+dclear the line · exit on an empty one

Input is multi-line and soft-wraps, and pasting is bracketed: a pasted block arrives as one message and one history entry, however many lines it has. /help lists these in-session.

Configuration

One TOML file: ~/.caravan/config.toml (or the path in CARAVAN_CONFIG). caravan init writes it (mode 0600). Edit it from any surface — no shell required:

caravan config set permissions ask     # CLI
/config set permissions ask            # REPL (and /key <provider> for keys)
# web UI: ⚙ settings panel

Editing the file

Every surface writes through the same schema, so a setting knows what it accepts before anything reaches disk:

caravan config keys                    # every setting, its value, what it accepts
caravan config set max_turns 40
caravan config unset base_url          # fall back to the provider default
caravan config check                   # validate the file against the schema
caravan config edit                    # $EDITOR, restored if the result won’t parse

A value that a setting cannot accept is refused rather than stored:

$ caravan config set model 3
✓ model = "3"                          # free text stays text
$ caravan config set permisions ask
Error: unknown setting 'permisions' — did you mean 'permissions'?
$ caravan config set max_turns 99999
Error: max_turns must be between 1 and 1000 (got 99999)

Writes are surgical: your comments, key order, and table layout survive an edit untouched — including /subagents add and /subagents remove, which append and delete [[subagents]] blocks in place. The previous contents are kept as config.toml.bak, and reading the config never modifies it.

/config with no arguments opens the settings as a list: arrow to one, press Enter, and pick a value (or type one, for the free-text settings). /config show keeps the old summary of the live session.

If a setting appears to do nothing, an environment variable is usually overriding it — config set, config get, and caravan doctor all say so when that is the case.

Diagnostics

caravan doctor checks the config file, the provider and its key, the transcript directory, subagents and MCP servers — and then offers to fix what it can: arrow to a failing check, press Enter, and the change is applied and the suite re-run.

caravan doctor           # report, then offer fixes on a terminal
caravan doctor --fix     # apply every fix that needs no input, then re-check
caravan doctor --json    # one JSON object, for CI

It exits non-zero when any check fails, so it can gate a script. /doctor runs the same checks inside a session, against that session's own connection.

Resolution order

  1. CLI flags — -p/--provider, -m/--model, --base-url, -s/--system;
  2. Environment — CARAVAN_PROVIDER, CARAVAN_MODEL, CARAVAN_BASE_URL, CARAVAN_STREAM, CARAVAN_MAX_TURNS, CARAVAN_PERMISSIONS, CARAVAN_TRANSCRIPT, CARAVAN_NUDGE, CARAVAN_SUBAGENTS, CARAVAN_STRICT_MODE, CARAVAN_SPINNER, CARAVAN_TLS_INSECURE, provider key vars (ANTHROPIC_API_KEY, …);
  3. TOML root keys, then the same keys inside [orchestrator];
  4. Registry / module defaults.

Core keys

provider    = "anthropic"          # see `caravan providers`
model       = "claude-sonnet-5"    # omit → provider default
# base_url  = "http://my-gateway:8000/v1"
system      = "You are a concise research assistant."   # appended to the shipped default
# system_replace = true  # make `system` replace the shipped default entirely

stream      = true       # stream tokens as they arrive
max_turns   = 15         # agent turn budget (default 24)
nudge       = true       # budget-awareness nudges in agent loops
permissions = "auto"     # auto | ask | readonly
provider_retry = "medium"  # provider error retry aggression: off | low | medium | high
# provider_retry_base_delay = 0.5  # base backoff seconds (exponential, cap 30s)
transcript  = true       # JSONL session logs in ~/.caravan/logs/
strict_mode = 0          # bash tool: 0 permissive, 1 single-command, 2 hidden
enable_subagents = true  # offer delegate when [[subagents]] exist

tool_call_mode = "auto"  # tool-call recognition: auto | native | text
require_finish = true    # agent runs complete only via the finish tool
tool_profile   = "auto"  # tool surface: auto (capability-driven) | core | full
# summarize_model = "…"  # cheap model for compaction summaries

Model capabilities

Unknown models get conservative defaults (8k context window, tool calling treated as unreliable). Override per model-name pattern (case-insensitive substring match); every field is optional:

[capabilities."my-local-model"]
context_window = 32768
tool_calling = "native"        # native | flaky | text
streaming_tool_calls = true
cache = "automatic"            # none | automatic | explicit
requests_per_minute = 20

The capability table drives the compaction threshold, the text tool-call fallback, the tool profile, and the system-prompt layers — see Getting Started for Free.

caravan config keys (or /config keys) lists every editable key with its current value, what it accepts, and when a change takes effect. It is generated from the same schema that validates writes, so it can never drift from what the code actually reads.

Provider retries

When a provider call fails transiently (HTTP 5xx, 429 rate limits, dropped connections), Caravan can retry it automatically instead of interrupting the turn. provider_retry controls the aggression:

ModeRetriesRetries on
off0never
low15xx + connection failures
medium (default)35xx + 429 + connection failures
highunlimitedevery HTTP status, including deterministic 4xx

When the server supplies Retry-After (or an x-ratelimit-reset-* header), that wait is honoured instead — clamped to 120s — which is what makes 429-heavy free tiers survivable. Otherwise backoff is exponential (provider_retry_base_delay, default 0.5s: 0.5s → 1s → 2s … capped at 30s). Streaming responses are only retried before the first token reaches your terminal, so output is never duplicated. Every retry is announced in the transcript as a provider_retry event. Env overrides: CARAVAN_PROVIDER_RETRY, CARAVAN_PROVIDER_RETRY_BASE_DELAY.

API keys

[api_keys]
anthropic = "sk-ant-..."
groq      = "gsk_..."

Environment variables take precedence and are the recommended home for secrets; the [api_keys] table is for hosts where a private 0600 file beats env plumbing (cron, HPC batch scripts).

Subagents

See Subagents for [[subagents]] and [providers.*] tables.

Spinner

[spinner]
enabled = true    # auto-disabled when stderr is not a TTY
verbose = false
thinking = ["Thinking", "Pondering", "Mulling"]   # arrays pick at random

MCP servers

[[mcp.servers]]
name      = "filesystem"
transport = "stdio"
command   = "npx"
args      = ["-y", "@modelcontextprotocol/server-filesystem", "/home/you/ws"]

Plugins

Caravan's tool composition runs on the plugin runtime (see Plugins). With no [[plugins]] table the default composition applies: the built-in tools plus one MCP mount per [[mcp.servers]] entry — existing configs behave exactly as before. Declare [[plugins]] entries to take control:

[[plugins]]                      # built-in tools, minus bash
plugin  = "tools.builtin"
exclude = ["bash"]

[[plugins]]                      # an MCP server as a plugin
id      = "fs"
plugin  = "tools.mcp"
name    = "filesystem"
command = "npx"
args    = ["-y", "@modelcontextprotocol/server-filesystem", "/home/you/ws"]
  • plugin names a builder (tools.builtin, tools.mcp, or one an embedding application registered); id defaults to it.
  • Entries merge over the defaults by id — redeclaring tools.builtin with enabled = false switches the default off.
  • An optional realm = "<name>" field sandboxes the entry's tools into an isolated realm that only [[subagents]] workers declaring the same realm can see — details in Subagents → Sandbox realms.
  • /plugins in the REPL shows each entry's lifecycle state; /plugins enable|disable <id> toggles one for the session.

Diagnostics

caravan doctor          # validity, key presence, reachability, subagents
caravan config show     # print the active file
caravan config path     # where it lives

A fully annotated example lives at docs/example_config.toml.

Providers

Caravan speaks the OpenAI chat-completions dialect to every backend through one engine (CaravanProviders.Openai_compatible), configured by a single data table (CaravanProviders.Registry). Run caravan providers to see this table live, with key status for your environment.

NameKindKey env varDefault model
ollamalocal—llama3.2
llama_cpplocal—default (whatever is loaded)
vllmlocal—default
lmstudiolocal—default
openaicloudOPENAI_API_KEYgpt-4o-mini
anthropiccloudANTHROPIC_API_KEYclaude-sonnet-5
groqcloudGROQ_API_KEYllama-3.3-70b-versatile
openroutercloudOPENROUTER_API_KEYmeta-llama/llama-3.3-70b-instruct
togethercloudTOGETHER_API_KEYmeta-llama/Llama-3.3-70B-Instruct-Turbo
deepseekcloudDEEPSEEK_API_KEYdeepseek-chat
mistralcloudMISTRAL_API_KEYmistral-small-latest
geminicloudGEMINI_API_KEYgemini-2.0-flash
xaicloudXAI_API_KEYgrok-3-mini
cerebrascloudCEREBRAS_API_KEYllama-3.3-70b
github_modelscloudGITHUB_TOKENopenai/gpt-4o-mini
nvidiacloudNVIDIA_API_KEYmeta/llama-3.3-70b-instruct

Model names drift faster than any table: caravan models (live, per provider) is authoritative, and the free-tier guide covers the zero-cost entries in detail.

Key resolution order

  1. The provider's environment variable (ANTHROPIC_API_KEY, …) — preferred;
  2. [api_keys] <provider> = "…" in ~/.caravan/config.toml (file is 0600);
  3. legacy openai_api_key top-level key (openai only).

A model for every weight class

caravan providers --ladder prints a curated pick per size:

ClassSuggestionNote
tiny ~1Bollama / llama3.2:1bruns on a laptop CPU
small ~4Bollama / qwen3:4bfast local reasoning
medium ~20Bollama / gpt-oss:20bstrong local, ~16 GB
large ~70Bgroq / llama-3.3-70b-versatileopen weights, hosted fast
frontieranthropic / claude-sonnet-5, openai / gpt-4o, gemini / gemini-2.5-pro

Notes and caveats

  • Anthropic and Gemini are reached through their official OpenAI-compatible endpoints, so tool calling and streaming work through the same code path as everyone else. A few provider-specific parameters (e.g. Anthropic's fine-grained thinking controls) are not exposed through that compatibility surface; if you need them, add a bespoke provider module — the PROVIDER signature is four functions.
  • Any other OpenAI-compatible server (text-generation-inference, llamafile, a lab-internal gateway): use --base-url (or base_url in the config) together with -p openai, plus OPENAI_API_KEY if the gateway wants auth.
  • TLS: certificates are verified against the system CA store, with hostname checking. CARAVAN_TLS_INSECURE=1 disables verification for self-signed lab endpoints (a warning is printed).

Getting Started for Free

You can run a working Caravan agent without paying anyone. This page lists the zero-cost routes, roughly in order of "least setup first", and the settings that make cheap models behave.

Free routes

Local models (no key, no meter)

ollama pull llama3.2        # or qwen3:4b, gpt-oss:20b …
caravan -p ollama -m llama3.2

Anything Ollama, llama.cpp, vLLM, or LM Studio can load works, on your hardware, with no rate limits. Small local models are the weakest models on this page but the strongest deal: unlimited requests.

GitHub Models — free with a token you already have

Any GitHub account token works (GITHUB_TOKEN), rate-limited but free:

export GITHUB_TOKEN=ghp_…
caravan -p github_models -m openai/gpt-4o-mini

OpenRouter :free models

One key, hundreds of models, and every model with a :free suffix costs nothing (about 20 requests/minute):

export OPENROUTER_API_KEY=sk-or-…
caravan -p openrouter -m "meta-llama/llama-3.3-70b-instruct:free"

Browse the current list at openrouter.ai/models.

Groq, Cerebras, NVIDIA NIM, Gemini

All four offer free tiers or free development credits with a signup:

ProviderSign up atNotes
groqconsole.groq.comvery fast 70B-class open weights
cerebrascloud.cerebras.aivery fast, generous daily quota
nvidiabuild.nvidia.comfree API credits, wide catalogue
geminiaistudio.google.comfree tier on flash models

Settings that matter on a free tier

Caravan derives most of this automatically from its capability table (see below), but these are the levers:

# ~/.caravan/config.toml

# Requests are the scarce resource. Rate-limit hits (429) honour the
# server's Retry-After automatically; "high" retries hardest.
provider_retry = "medium"

# Recognise tool calls that small models emit as text (default "auto").
tool_call_mode = "auto"

# Reduced tool surface for small-context models (default "auto").
tool_profile = "auto"

# Route compaction summaries to a cheap/free model so they don't burn
# the working model's quota.
summarize_model = "meta-llama/llama-3.3-70b-instruct:free"

Telling Caravan about your model

Unknown models get conservative defaults (8k context, tool calling treated as unreliable). If you know better, say so:

[capabilities."my-local-model"]
context_window = 32768
tool_calling = "native"     # native | flaky | text
requests_per_minute = 20

The pattern is matched as a case-insensitive substring of the model name; every field is optional and patches over the built-in table.

What the harness already does for you

  • A default system prompt and environment preamble, so small models don't waste metered requests rediscovering the working directory.
  • Text tool-call recovery: a model that prints {"tool": "ls", "arguments": {}} instead of using native tool calls still works (auditable via tool_call_fallback trace events).
  • Token-aware compaction sized to the model's context window, with a free structural tier before any summarisation call is spent.
  • Byte-stable request prefixes, so providers with automatic prompt caching (DeepSeek, OpenAI) bill cached rates for the shared history.
  • An enforceable turn budget (max_turns, default 24) so a confused model cannot silently drain a daily quota.

Agents, Permissions & Transcripts

The agent loop

/agent and caravan agent run a ReAct-style loop: the model calls tools, sees results, and signals completion with the finish tool. The loop is budgeted (max_turns) and nudged: at the halfway point and near exhaustion, Caravan reminds the model of the task and remaining turns — this measurably reduces wandering in long runs. Disable with nudge = false.

Tools

Run /tools to see the active set (✎ marks mutating tools):

bash · read_file / write_file · ls / grep / sed / touch / mkdir · web_fetch / web_search · lisp · summarize · finish · delegate (when subagents are configured) · plus any MCP-server tools from the config.

Permissions

permissions = "auto"   # auto | ask | readonly
ModeBehavior
autoall tools run (default)
askinteractive y/n/always prompt before each mutating tool
readonlymutating tools denied outright — audit-safe runs

Mutating: bash, write_file, sed, touch, mkdir, delegate. Switch live with /permissions ask, per-run with CARAVAN_PERMISSIONS=readonly. In the web UI, ask degrades to deny (there is no prompt surface).

Transcripts

With transcript = true (default) every session appends structured events to ~/.caravan/logs/session-<timestamp>-<pid>.jsonl:

{"ts":1786349217.4,"event":"tool_call_start","name":"bash","args":"{\"command\":\"dune test\"}"}
{"ts":1786349219.1,"event":"tool_call_end","name":"bash","output":"…","duration_s":1.7}
{"ts":1786349226.0,"event":"task_finished","summary":"All tests pass after the fix."}

caravan agent --json includes the transcript path, so pipelines can archive exactly what their agent did. Under the hood this is Caravan.Trace — one event stream feeding the terminal renderer, the JSONL sink, and anything you attach (see Library Guide).

Memory

Sessions keep a sliding window (/memory <n>, 0 = unlimited) and can compact on demand (/summarise) or automatically when the window overflows. Summary and hierarchical memories are available to library users, plus Redis for shared multi-process context.

Subagents

Declare workers in the config and the orchestrator model gets a delegate tool — no OCaml required. Each worker starts cold (no history bleed), runs its own tool loop on its own provider/model, and returns a compact result.

Declaring workers

# Optional master switch (default: on when [[subagents]] tables exist)
enable_subagents = true

# A worker on a lab endpoint, declared explicitly:
[providers.local_qwen]
base_url    = "http://127.0.0.1:8080/v1"
api_key_env = "QWEN_KEY"          # optional

[[subagents]]
name          = "coder"
provider      = "local_qwen"      # a [providers.*] table, or a registry name
model         = "qwen3:8b"
tools         = ["bash", "write_file", "read_file"]
system_prompt = "You write minimal, correct OCaml."
temperature   = 0.2               # optional
max_tokens    = 2048              # optional

# A worker on a registry provider — just name it:
[[subagents]]
name          = "reviewer"
provider      = "anthropic"
model         = "claude-haiku-4-5"
tools         = []                # finish is always added
system_prompt = "You review diffs ruthlessly but concisely."

Using them

  • /subagents shows the roster with provider/key health;
  • the model calls delegate {"subagent": "coder", "task": "…"} — the task must be self-contained (workers have no memory of the parent conversation);
  • results stream back into the orchestrator's context.

Sandbox realms

A worker can carry a plugin-toolset sandbox (see Plugins):

[[subagents]]
name  = "researcher"
provider = "ollama"
model = "qwen3:8b"
tools = ["read_file", "web_search"]   # explicit whitelist, as before
realm = "research"                    # + everything plugins put here

[[plugins]]
id     = "research-mcp"
plugin = "tools.mcp"
realm  = "research"                   # this server's tools go ONLY to
command = "npx"                       # workers with realm = "research"
args   = ["-y", "@modelcontextprotocol/server-arxiv"]

Semantics:

  • tools stays the explicit whitelist against the shared toolset — nothing changes for existing configs;
  • realm adds whatever plugins registered into the named realm, resolved at delegation time — plugins loading, unloading, or being toggled between delegations take effect on the next delegate call, with no restart. On a name collision the whitelisted tool wins;
  • realm tools never appear in the orchestrator's toolset, and workers without the realm never see them — the same key, isolated bindings (the paper's coeffect isolation);
  • /subagents shows each worker's realm and its current sandbox size.

Governance

  • delegate is a mutating tool: ask prompts, readonly denies;
  • every delegation and each worker's own tool calls flow through the transcript — the audit trail covers the whole tree;
  • misconfigured entries (unknown provider, missing tool) degrade to startup warnings, never crashes; caravan doctor validates the roster;
  • toggle without editing files: /config set enable_subagents false or CARAVAN_SUBAGENTS=0.

Programmatic swarms

For dynamic spawning (models designing their own workers), see the two swarm examples in examples/ and Caravan.Subagent in the API docs.

Slip — the micro-LISP

Slip is Caravan's embedded symbolic engine: a tiny, total (step-capped, always terminates), sandboxed (no IO, no clock, no randomness) LISP over the JSON universe. Models use it through the lisp tool to do exact arithmetic, counting, filtering, and data reshaping instead of approximating in prose; humans poke it with /lisp.

❯ /lisp (mean (map (lambda (r) (get "t" r)) data))

Why

Long agentic runs die on small arithmetic and bookkeeping mistakes. Handing a model — 2B or 200B — an exact calculator with its full manual in the tool description removes a whole failure class. Homoiconicity (code = lists) plays to model strengths: programs are easy to generate, inspect, and even transform with read / show / eval.

The language on one page

; forms
(if cond then else)  (let ((x 1) (y 2)) body)  (define name expr)
(lambda (x) body)    (do e1 e2)                (quote x)  or  'x

; math / compare        (+ - * / mod abs min max round sum mean)
(+ 1 2 3)        ; 6    (= != < > <= >= not and or)
(sum (range 1 101))     ; 5050

; lists
(list 1 2)  (len xs)  (first xs)  (last xs)  (nth 0 xs)  (rest xs)
(append a b)  (reverse xs)  (sort xs)  (range 0 10)
(map f xs)  (filter f xs)  (reduce f init xs)
(map upper (list "a" "b"))            ; builtins are first-class

; records & tables (JSON objects / arrays of objects)
(get "name" row)  (keys row)  (put "k" v row)
(select "name" "age" rows)  (where "role" "admin" rows)  (sort-by "age" rows)

; strings
(str "n=" 3)  (upper s)  (lower s)  (contains s sub)
(split "a,b" ",")  (join xs ", ")

; types & JSON
(number? x) (string? x) (list? x) (record? x) (null? x)
(parse-json "[1,2]")  (to-json x)

; code as data
(read "(+ 1 2)")   ; → the list (+ 1 2)
(show '(+ 1 2))    ; → "(+ 1 2)"
(eval '(+ 1 2))    ; → 3

Recursion works and is safe — the evaluator burns a step budget (default 100 000) and returns a clean error instead of hanging:

(define fact (lambda (n) (if (<= n 1) 1 (* n (fact (- n 1))))))
(fact 10)   ; 3628800

The lisp tool

Input: {"program": "...", "data": <any JSON>} — data is bound to the symbol data inside the program. Native data operations (sum, sort, where, …) cost one step regardless of size, so million-element folds are fine; only interpreted steps (lambda applications) burn budget.

From OCaml: Caravan.Lisp.run ?max_steps ?data src → (Value.t, string) result.

Web UI

caravan web [--port 8787]

A single, fully embedded page — no assets on disk, no JS toolchain, no CDN dependencies for the app itself — served on 127.0.0.1 only. It is a personal cockpit, not a deployment target.

What it does

  • Chat with the active model, or tick agent to run autonomous tasks; every reply shows the tool calls that produced it plus token usage;
  • ⚙ Settings: edit every config key and paste API keys from the browser — the no-shell configuration path. Key values are write-only (the API reports only set/unset) and land in the same 0600 TOML file;
  • live provider/model/token counters in the header.

JSON API

EndpointPurpose
GET /api/stateprovider, model, token totals, permission mode
POST /api/chat {"message": "…"}one chat turn
POST /api/agent {"task": "…"}autonomous run; response includes a tools audit trail
GET /api/configeditable settings + per-provider key presence (never key values)
POST /api/config {"key","value"}edit a whitelisted setting
POST /api/key {"provider","key"}store an API key

Security posture

  • binds loopback only; no auth layer — do not port-forward it;
  • POST /api/config is whitelisted to the documented settings keys; arbitrary TOML paths are rejected;
  • permission modes apply to web runs too; ask degrades to deny since no prompt is possible.

Library Guide

Caravan is three findlib libraries: Caravan (core), CaravanProviders, CaravanTools. Full signatures live in the API reference.

A typed pipeline in ten lines

open Caravan
open Caravan.Chain

let fact_chain net provider =
  prompt_template "List 3 facts about {{topic}}."
  |>> llm net provider
  |>> parse Parser.numbered_list

let () = Eio_main.run (fun env ->
  let provider = CaravanProviders.Ollama.make_provider ~model:"llama3.2" () in
  match run (fact_chain env#net provider) [("topic", "OCaml")] with
  | Ok facts -> List.iter print_endline facts
  | Error e  -> prerr_endline e)

|>> is Result-bind: every stage is 'a -> ('b, string) result. Also available: parallel, retry ~n, Kleisli >=>.

Sessions and agents

let sess = Session.create ~tools:CaravanTools.All_tools.all_tools model provider in
let sess = Session.set_system sess "Be terse." in
match Agent.run env#net env#clock sess "count the .ml files here" with
| Ok (_sess, result) -> print_endline result.value.content
| Error e -> prerr_endline e

Wrap calls in Effects.with_net env#net so network tools reuse your event loop, and in Effects.run_with_effects ~permission_policy:(Permission.policy_of_mode "ask") to govern tools.

Writing a tool

A tool is a module satisfying Caravan.Tool.TOOL — typed input/output, a JSON schema, and an execute. Drop the file in lib/tools/; the build-time generator registers any file that defines module <Capitalized-filename>:

(* lib/tools/word_count.ml *)
open Caravan.Tool

module Word_count = struct
  let name = "word_count"
  let aliases = ["wc"]
  let description = "Counts words in a string."
  type input = string
  type output = int

  let json_schema () = `Assoc [
    "type", `String "object";
    "properties", `Assoc ["text", `Assoc ["type", `String "string"]];
    "required", `List [`String "text"] ]

  let parse_args json =
    try Ok Yojson.Safe.Util.(json |> member "text" |> to_string)
    with _ -> Error "expected {\"text\": …}"

  let format_output n = string_of_int n

  type _ Effect.t += Exec : input -> output Effect.t
  let execute text =
    String.split_on_char ' ' text |> List.filter (( <> ) "") |> List.length
end

If the tool mutates state, add its name to Permission.mutating_tools.

Writing a provider

Any OpenAI-compatible endpoint needs no code — add a registry entry (one record in lib/providers/registry.ml) or just use base_url. For a genuinely different API, implement Caravan.Provider.PROVIDER (four functions: name, complete, stream, list_models) and pack it with Provider ((module M), cfg).

Observing everything: Trace

Trace.add_sink (function
  | Trace.Tool_call_end { name; duration; _ } ->
    Printf.eprintf "%s took %.1fs\n" name duration
  | _ -> ());

(* or capture scoped: *)
let (result, events) =
  let acc = ref [] in
  let r = Trace.with_sink (fun e -> acc := e :: !acc) run_my_agent in
  (r, List.rev !acc)

Trace.open_transcript ~dir attaches the JSONL sink the CLI uses.

Memory back-ends

Memory.Ring (sliding window), Memory.Summary (compact-on-demand), Memory.Hierarchical (rolling summaries), Redis_store (shared, multi-process). Sessions accept any MEMORY implementation via the packed existential.

Plugins — Spatiotemporal Composability

Caravan.Plugin is a dynamic-composition runtime: components that can be loaded, unloaded, and rewired while the system runs, with the runtime — not programmer discipline — guaranteeing that removal reverts every side effect and that dependencies stay consistently wired.

It is an OCaml implementation of the model in "A Programming Paradigm for Spatiotemporal Composability" (Shi, Zhang, Cui — Peking University / DeepSeek-AI, 2026), the formal foundation behind the Cordis framework and the DeepSeek agent harness (dsh). The paper identifies two orthogonal guarantees:

  • Temporal composability — every mutation a component makes carries an inverse the runtime tracks; unloading replays the inverses in LIFO order, so the environment is recovered exactly.
  • Spatial composability — components declare what they inject (read from the environment) and provide (write to it); the runtime resolves the declarations reactively, activating a component when its dependencies are present and deactivating it when they go away.

Why this matters for an agent harness: a harness that composes tool suites, providers, MCP servers, and memory backends at runtime — or that one day installs components an agent generated for itself — needs to remove a faulty component without restarting, and needs dependents to react when a component is replaced. That is exactly what this runtime makes structural.

Quick start

open Caravan

(* A typed service key. *)
let db_key : Db.t Plugin.Key.t = Plugin.Key.create ~name:"database" ()

(* A provider component: declares what it may provide. *)
let db_plugin =
  Plugin.component ~name:"db" ~provide:[ Plugin.Key.Ex db_key ]
    (fun ctx ->
      let db = Db.connect () in
      ignore (Plugin.provide ctx db_key db);
      ignore (Plugin.on_dispose ctx (fun () -> Db.close db)))

(* A consumer component: declares what it injects. It will not run
   until the database is provided, and is deactivated (cleanly, with
   the db still readable during its teardown) when the db goes away. *)
let api_plugin =
  Plugin.component ~name:"api" ~inject:[ Plugin.Key.Ex db_key ]
    (fun ctx ->
      let db = Plugin.get ctx db_key in
      let server = Server.start db in
      ignore (Plugin.on_dispose ctx (fun () -> Server.stop server)))

let () =
  let ctx = Plugin.make () in
  let _api = Plugin.use ctx api_plugin in   (* pending: no db yet *)
  let db = Plugin.use ctx db_plugin in      (* both activate *)
  Plugin.dispose db                          (* api deactivates first,
                                                then db tears down *)

Run the offline demo: dune exec examples/plugin_system/plugin_system.exe.

The model in five pieces

Paper conceptAPIWhat it does
revertible effectPlugin.track ctx runrun a setup step now, register its inverse; LIFO replay on unload
coeffect provisionPlugin.provide ctx key vbind a typed service; tracked, withdrawal ordered after dependents drain
coeffect accessPlugin.get ctx keyread a service, checked against the component's declarations
componentPlugin.component ~inject ~provide bodydeclarations + effectful body
fiberPlugin.use ctx compone live instantiation with a lifecycle: Pending → Loading → Active → Unloading → … (or Failed)

Everything a component does to the outside world goes through its context: track, provide, on (event listeners), use (child components), Toolset.register. Because each is tracked, teardown is derived from loading — components rarely need explicit cleanup code beyond on_dispose for resources the runtime cannot see.

The discipline (and its error messages)

Reads and writes are checked against declarations at the point of use:

ViolationException
get on a key the component never declaredUndeclared_access
get while the declaring fiber is not committed to a providerInactive_access
get at root with no bindingUnprovided
provide on a key outside the component's provide listUndeclared_provision
second provider for a key in the same realmDuplicate_provider
creating effects on a non-loading, non-active contextInactive_context

A component body that raises (including any of the above) marks its fiber Failed and rolls back the effects it had installed; siblings are untouched. Plugin.restart retries a failed fiber explicitly — the runtime deliberately does not retry a body that has proven unsound against an unchanged environment (paper §4.3.4).

Isolation and interception

Plugin.isolate ctx (Key.Ex key) derives a context in which key resolves in a separate realm — a private one by default, or a shared named one with ~realm:"tenant1". Components instantiated on the derived context inherit it: the same key, different binding. Use it for sandboxed subagents, per-session tool tables, or test doubles.

Plugin.intercept ctx (Key.Ex key) json attaches metadata to accesses of key through the derived context; a service implementation can consult it with Plugin.interception to adjust behaviour per consumer (timeouts, rate limits, permissions) without the consumer changing.

Events

Plugin.Event.create, Plugin.on, and Plugin.emit form a typed event bus whose listener registrations are tracked effects — a plugin's listeners disappear with the plugin.

The toolset service

Plugin.Toolset bridges the runtime to Caravan's agents: a shared registry of Tool.packed_tools that plugins extend revertibly.

let ctx = Plugin.make () in
ignore (Plugin.use ctx Plugin.Toolset.provider);
let pack =
  Plugin.component ~name:"fs-tools"
    ~inject:[ Plugin.Key.Ex Plugin.Toolset.key ]
    (fun ctx ->
      ignore (Plugin.Toolset.register ctx
        (Tool.Tool (module CaravanTools.Read_file.Read_file))))
in
ignore (Plugin.use ctx pack);
(* hand the live tool list to an agent *)
let tools = Plugin.Toolset.snapshot ctx in

snapshot gives an ordinary Tool.packed_tool list for Session.create / Agent.run; call it per run to pick up whatever the plugin set currently provides.

Declarative reconciliation

Plugin.Reconcile turns an entry list — id, enabled flag, JSON config, component builder — into the corresponding set of running fibers, and incrementally reconciles on every apply: new entries instantiate, removed ones dispose, and a changed config or flag rebuilds just that entry. This is the loader pattern of the paper's §5.2.1: an orchestrator edits a description; the runtime performs the minimal transitions.

The harness runs on it

The CLI's tool composition is itself plugin-hosted (Plugin_host, the policy layer over the runtime):

  • Built-in tools and MCP servers are plugin fibers. Plugin_host keeps a registry of named builders (tools.builtin, tools.mcp, plus any an embedding application registers) and reconciles [[plugins]] config entries into running fibers — with the classic composition (built-ins + [[mcp.servers]]) synthesized as the default when the table is absent. See Configuration → Plugins.
  • MCP mounts are revertible: disposing an tools.mcp fiber closes the server process and withdraws its tools; a failed connection is a Failed fiber, visible in /plugins, that never disturbs siblings.
  • /plugins lists each entry with its live lifecycle state; /plugins enable|disable <id> reconciles one entry and refreshes the session's toolset in place.
  • Subagent workers can be sandboxed — a [[plugins]] entry with a realm = "<name>" field registers its tools into an isolated toolset realm instead of the shared one (the entry is instantiated with Toolset.key isolated — the isolate annotation of the paper's loader entries, Def. 74), and a [[subagents]] worker declaring the same realm resolves those tools at delegation time. See Subagents → Sandbox realms.
  • The active provider is a service (Plugin.Services.provider). The CLI provides it at session setup and re-provides it on /provider and /model switches, so a plugin that injects it reloads against the new provider automatically.
  • Lifecycles are auditable: every fiber transition is a Trace.Plugin_transition event (verbose mode prints them; the JSONL transcript always records them), and run failures at the REPL/agent boundaries are recorded as Trace.Run_error events — failed sessions leave transcripts too.

Scope and honest limits

  • Synchronous core. Transitions are the paper's base calculus + failure: atomic and run-to-completion. The asynchrony/inertia layer (paper §4.3.3) is not implemented; if a body blocks, use blocks. An Eio-fibered lifecycle is a natural future layer.
  • No hot code loading. OCaml links statically; components are OCaml values, so "replacing a module" means rebuilding a fiber from a new component value (Reconcile covers the config-driven case). The paper's HMR chapter applies to its JavaScript host, not here.
  • Single-source is sticky. A second provider for an occupied key fails at activation and stays Failed until an explicit restart, rather than being refused at instantiation.
  • Interception carries context metadata only — the paper's per-key metadata monoids and component-declared metadata are simplified to JSON with nearest-wins shallow merge.

See docs/COMPOSABILITY_NOTES.md for the running friction log of this subsystem.

Architecture

Caravan separates what happened (the event stream), what to do (the agent loop), and how to talk (providers/tools) — so every layer can be swapped or observed without touching the others.

flowchart TB
    Entry["cli/cli.ml<br/>(CLI: repl · agent · web)"]

    subgraph Frontends ["Front-ends (cli/)"]
        Editor["editor.ml · tty.ml<br/>(multi-line editor,<br/>palette, history, paste)"]
        Picker["picker.ml · commands.ml<br/>(select/form widgets,<br/>the command table)"]
        Render["render.ml<br/>(Trace → terminal)"]
        WebUI["web.ml<br/>(localhost cockpit)"]
        SubW["subagents.ml<br/>(config → delegate tool)"]
    end

    subgraph Core ["The Brain (lib/)"]
        Agent["agent.ml<br/>(loop + budget nudges)"]
        Session["session.ml<br/>(history, tool dispatch)"]
        Memory["memory.ml<br/>(ring / summary / hierarchical)"]
        Trace["trace.ml<br/>(event stream + JSONL)"]
        Lisp["lisp.ml<br/>(Slip micro-LISP)"]
        Tls["tls.ml<br/>(verified HTTPS)"]
    end

    subgraph Backends ["Pluggable Backends"]
        direction LR
        Registry["providers/registry.ml<br/>(13 backends, one engine)"]
        Tools["tools/<br/>(fs · shell · web · lisp · delegate)"]
    end

    Config["config.ml ⇄ ~/.caravan/config.toml<br/>(one 0600 file; CLI/REPL/web editors)"]

    Entry --> Editor
    Entry --> WebUI
    Editor --> Session
    Render -.listens.-> Trace
    WebUI -.listens.-> Trace
    Agent <--> Session
    Session --- Memory
    Session -.emits.-> Trace
    Session --> Tools
    Tools --> Lisp
    Agent ==> Registry
    Registry --> Tls
    SubW --> Tools
    Config -.-> Agent
    Config -.-> Registry
    Config -.-> SubW

Load-bearing decisions

  • Events, not prints. lib/ never writes to the terminal; it emits Trace events. The CLI renderer, the JSONL transcript, and the web audit trail are just sinks. Auditability falls out for free.
  • One wire dialect. Every provider speaks OpenAI chat-completions (Anthropic/Gemini via their official compat endpoints), so one engine (Openai_compatible) plus a data table (Registry) covers 13 backends. Exotic APIs implement the 4-function PROVIDER signature.
  • Effects for capabilities. Tools request the ambient network with the Get_net effect; permission checks are the Ask_permission effect. Front-ends install handlers once; the library stays pure of policy.
  • Packed existentials everywhere. Tools, providers, and memories are first-class modules packed with their state — heterogeneous lists with full type safety at the boundaries.
  • Totality where models roam. Slip is step-capped; agent loops are turn-budgeted and nudged; the bash tool reports exit codes honestly. Nothing a model does can hang the harness.

Module index

ModuleRole
Caravan.Typesmessages, roles, results; wire vs export JSON
Caravan.Sessionstateful conversations, tool execution
Caravan.Agentautonomous loops, budgets, nudges
Caravan.Traceevent stream, JSONL transcripts
Caravan.Tool / Effectstyped tools, effect dispatch, permissions
Caravan.Lispthe Slip engine
Caravan.Tlsthe single certificate-verifying HTTPS path
Caravan.Memory / Redis_storecontext strategies
Caravan.Chain / Parser / Template / Promptthe pipeline DSL
CaravanProviders.Registrythe provider table
CaravanTools.*the tool set (auto-registered at build time)

Development

Build & test

opam install . --deps-only --with-test -y
dune build          # zero warnings expected
dune test           # all offline — mock providers, no network
dune build @doc     # odoc API docs (requires odoc)

CI

  • .github/workflows/ci.yml — build + test on OCaml 5.2 / 5.3, plus a check that Caravan.opam stays in sync with dune-project;
  • .github/workflows/deploy.yml — on push to main, builds this book (mdBook) and the odoc API reference and publishes both to GitHub Pages (/ = book, /api/ = odoc).

Testing philosophy

Tests never touch the network: providers are mocked as first-class modules. For end-to-end verification during development we drive the real binary against a scripted OpenAI-compatible mock server and a PTY harness for the line editor — every fixed bug gets a regression test.

Repository map

bin/        CLI: main, line editor, Trace renderer, web UI, subagent wiring
lib/        core: types, session, agent, trace, tls, lisp, memory, …
lib/providers/  the OpenAI-compatible engine + registry
lib/tools/      tool modules (auto-registered at build time)
docs/       this book (src/) + example config + overhaul notes
examples/   single-endpoint config, two subagent-swarm programs
scripts/    installer, docs renderer
test/       inline test suite (ppx_expect / ppx_inline_test)

Extending safely

  • New tool → Library Guide; mark it mutating if it changes state.
  • New provider → registry entry, or a PROVIDER module for exotic APIs.
  • Anything user-visible → emit Trace events, never print from lib/.
  • New plugin / dynamic component → Plugins; go through the context (track, provide, on) so unloading stays revertible.
  • Keep the pain-point log honest: docs/COMPOSABILITY_NOTES.md records friction, divergences from the paper, and why decisions fell the way they did.

API Reference

Caravan provides three findlib libraries:

  • Caravan — Core framework, ReAct agents, typed chains, session state, event tracing, memory compaction, micro-LISP, and MCP client.
  • CaravanProviders — Pluggable backends (Ollama, OpenAI, llama.cpp, Groq, Anthropic, Gemini, DeepSeek, generic OpenAI-compatible endpoints).
  • CaravanTools — Built-in executable tools (Bash shell execution, file read/write, web fetch/search, subagent delegation, LISP runner).

Browsing the complete, auto-generated OCaml odoc HTML documentation is available at: Full HTML API Reference (odoc)


Core Library: Caravan

Caravan.Agent

Autonomous agentic loop (ReAct loop) execution engine.

type agent_config = {
  max_turns : int;
  continue_prompt : string;
  nudge : bool;
}

val default_config : unit -> agent_config

val run :
  ?config:agent_config ->
  ?on_turn:(Session.t -> chat_message result_with_meta -> unit) ->
  ?on_step:(Session.t -> unit) ->
  _ Eio.Net.t ->
  _ Eio.Time.clock ->
  Session.t ->
  string ->
  (Session.t * chat_message result_with_meta, string) result

val run_stream :
  ?config:agent_config ->
  ?on_turn:(Session.t -> chat_message result_with_meta -> unit) ->
  ?on_step:(Session.t -> unit) ->
  _ Eio.Net.t ->
  _ Eio.Time.clock ->
  Session.t ->
  string ->
  on_token:(string -> unit) ->
  (Session.t * chat_message result_with_meta, string) result

Caravan.Chain

Composable typed LLM processing pipelines using Result-bind (|>>).

type ('a, 'b) t = 'a -> ('b, string) result

val (|>>) : ('a, 'b) t -> ('b, 'c) t -> ('a, 'c) t
val run : ('a, 'b) t -> 'a -> ('b, string) result
val run_exn : ('a, 'b) t -> 'a -> 'b

val prompt_template : string -> (string * string) list -> (string, string) result
val prompt_messages : ?system:string -> string -> (string * string) list -> (chat_message list, string) result
val llm : _ Eio.Net.t -> Provider.packed_provider -> chat_message list -> (string, string) result
val llm_stream : _ Eio.Net.t -> Provider.packed_provider -> on_token:(string -> unit) -> chat_message list -> (string, string) result
val parse : 'a Parser.t -> string -> ('a, string) result

val parallel : Eio.Switch.t -> ('a, 'b) t list -> ('a, 'b list) t
val retry : n:int -> ('a, 'b) t -> 'a -> ('b, string) result

module Kleisli : sig
  val compose : ('a -> ('b, 'e) result) -> ('b -> ('c, 'e) result) -> 'a -> ('c, 'e) result
  val ( >=> ) : ('a -> ('b, 'e) result) -> ('b -> ('c, 'e) result) -> 'a -> ('c, 'e) result
end

Caravan.Session

Multi-turn session history, system prompt management, and state persistence.

type t

type config = {
  model               : string;
  system              : string option;
  options             : gen_options;
  memory_size         : int;
  max_tool_output_len : int option;
  auto_summarize      : bool;
} [@@deriving yojson]

val create : ?tools:Tool.packed_tool list -> string -> Provider.packed_provider -> t
val set_system : t -> string -> t
val add_user : t -> string -> t
val add_assistant : t -> string -> t
val turn_idx : t -> int
val history : t -> chat_message list

val export_json : t -> Yojson.Safe.t
val of_json : provider:Provider.packed_provider -> ?tools:Tool.packed_tool list -> Yojson.Safe.t -> (t, string) result
val save_checkpoint : ?path:string -> t -> (string, string) result
val load_checkpoint : provider:Provider.packed_provider -> ?tools:Tool.packed_tool list -> ?path:string -> unit -> (t, string) result

Caravan.Trace

Auditable event stream for LLM completions, tool calls, nudges, and summarization.

type event =
  | Session_start of { provider : string; model : string }
  | Model_call_start of { prompt_len : int }
  | Model_call_end of { duration_s : float; usage : usage_stats option }
  | Tool_call_start of { name : string; args : string }
  | Tool_call_end of { name : string; output : string; duration_s : float }
  | Nudge of { content : string }
  | Task_finished of { summary : string }

type sink = event -> unit

val add_sink : sink -> unit
val with_sink : sink -> (unit -> 'a) -> 'a
val emit : event -> unit
val open_transcript : dir:string -> unit -> sink

Caravan.Types

Core message types, roles, results, and token usage records.

type role = System | User | Assistant | Tool

type tool_call = {
  id : string;
  name : string;
  args : string;
}

type chat_message = {
  role : role;
  content : string;
  name : string option;
  tool_calls : tool_call list option;
  tool_call_id : string option;
}

type usage_stats = {
  prompt_tokens : int;
  completion_tokens : int;
  total_tokens : int;
}

type 'a result_with_meta = {
  value : 'a;
  usage : usage_stats option;
  finish_reason : string option;
}

Caravan.Provider

Abstract provider interface and packed existential types.

module type PROVIDER = sig
  type config
  val name : string
  val complete : _ Eio.Net.t -> config -> chat_message list -> chat_message result_with_meta
  val stream : _ Eio.Net.t -> on_token:(string -> unit) -> config -> chat_message list -> chat_message result_with_meta
  val list_models : _ Eio.Net.t -> config -> (string list, string) result
end

type packed_provider = Provider : (module PROVIDER with type config = 'c) * 'c -> packed_provider

val name_of_packed : packed_provider -> string
val complete_packed : _ Eio.Net.t -> packed_provider -> chat_message list -> chat_message result_with_meta
val stream_packed : _ Eio.Net.t -> on_token:(string -> unit) -> packed_provider -> chat_message list -> chat_message result_with_meta

Caravan.Tool

First-class module tool interface and effect-based tool execution handler.

module type TOOL = sig
  val name : string
  val aliases : string list
  val description : string
  type input
  type output
  val json_schema : unit -> Yojson.Safe.t
  val parse_args : Yojson.Safe.t -> (input, string) result
  val format_output : output -> string
  type _ Effect.t += Exec : input -> output Effect.t
  val execute : input -> output
end

type packed_tool = Tool : (module TOOL) -> packed_tool

val name_of_packed : packed_tool -> string
val description_of_packed : packed_tool -> string
val schema_of_packed : packed_tool -> Yojson.Safe.t
val find_tool : packed_tool list -> string -> packed_tool option

Caravan.Subagent

Isolation and delegation of background sub-tasks to dedicated worker subagents.

type subagent_spec = {
  name : string;
  system_prompt : string;
  tools : Tool.packed_tool list;
  provider : Provider.packed_provider option;
  model : string option;
  max_turns : int option;
}

val delegate :
  _ Eio.Net.t ->
  _ Eio.Time.clock ->
  Session.t ->
  subagent_spec ->
  string ->
  (Session.t * chat_message result_with_meta, string) result

Caravan.Permission

Security permission policies governing tool execution.

type mode = Auto | Ask | Readonly

type policy = {
  mode : mode;
  prompt_user : string -> string -> bool;
}

val policy_of_mode : ?prompt_user:(string -> string -> bool) -> string -> policy
val is_mutating : string -> bool

Caravan.Memory

Context window management and history compaction.

module type MEMORY = sig
  type t
  val create : capacity:int -> t
  val add : t -> chat_message -> t
  val get : t -> chat_message list
  val clear : t -> t
end

type packed_memory = Memory : (module MEMORY with type t = 'm) * 'm -> packed_memory

module Ring : MEMORY
module Summary : MEMORY
module Hierarchical : MEMORY

Caravan.Lisp

Slip — Caravan's embedded micro-LISP interpreter for programmatic tool composition and evaluation.

type expr =
  | Symbol of string
  | String of string
  | Number of float
  | List of expr list
  | NativeFun of (expr list -> (expr, string) result)

val parse : string -> (expr list, string) result
val eval : env:(string, expr) Hashtbl.t -> expr -> (expr, string) result
val eval_string : env:(string, expr) Hashtbl.t -> string -> (expr, string) result

Caravan.Mcp

Model Context Protocol (MCP) tool integration and server connection registry.

type mcp_server_config = {
  name : string;
  transport : string;
  command : string;
  args : string list;
}

val load_mcp_tools : mcp_server_config list -> (Tool.packed_tool list, string) result

Provider Library: CaravanProviders

  • CaravanProviders.Ollama: Connects to local Ollama daemon (http://localhost:11434).
  • CaravanProviders.Openai: OpenAI API connector (GPT-4o, GPT-4o-mini, O3-mini).
  • CaravanProviders.Llama_cpp: Local llama.cpp HTTP server connector.
  • CaravanProviders.Openai_compatible: Generic OpenAI-compatible backend connector for vLLM, DeepSeek, Groq, Together, Mistral, OpenRouter, LM Studio, XAI, etc.
  • CaravanProviders.Registry: Global registry mapping provider names to instances and default models.

Tools Library: CaravanTools

  • CaravanTools.Bash: Executes shell commands with strict-mode safety options.
  • CaravanTools.Read_file / CaravanTools.Write_file: File I/O tools.
  • CaravanTools.Web_search / CaravanTools.Read_browser_page: Web research tools.
  • CaravanTools.Delegate: Subagent task delegation tool.
  • CaravanTools.Finish: Task completion tool.
  • CaravanTools.Lisp: Programmatic Slip LISP script runner tool.
  • CaravanTools.All_tools: Registry exporting all_tools : Tool.packed_tool list.