Ivy Router — routeur de modèles 100% use case perso : pré-filtre + classifieur SmolLM2 (Ollama) + sticky session + profil Lily. POC validé 65/68, QLoRA next.
Find a file
Ivy 558129f482 feat(cascade): le cloud est une décision, jamais un filet (décision Lily 12/09)
- les plans automatiques deviennent local-only : plus aucun palier cloud
  en position de repli (chat/tasks, dev-light, sysadmin, dev-heavy sans
  directCloud) ; un cloud ne peut plus être qu'en tête, là où §7.2 l'a décidé
- plan épuisé -> message dégradé explicite, jamais d'escalade silencieuse
- 97 lignes du dataset v2 + tests recalculés (518/518)
2026-09-12 09:19:59 +02:00
data feat(cascade): le cloud est une décision, jamais un filet (décision Lily 12/09) 2026-09-12 09:19:59 +02:00
docs docs: état des lieux v2 (12/09) + rapports de recherche et d'impact 9.4 2026-09-12 09:04:44 +02:00
legacy feat(scaffold): milestone 1 — typed config, keywords, journal, no-drift install 2026-09-07 14:09:27 +02:00
modelfiles chore(config): setup migration principal — workspace /home/ivy/.openclaw/workspace + localFallback Gemini explicite + Modelfile smollm2-cls-q4 (FROM adapté smollm2:1.7b) 2026-08-18 16:55:55 +02:00
scripts feat(dataset): milestone 9 — dataset v2 + non-regression 2026-09-08 13:14:45 +02:00
src feat(cascade): le cloud est une décision, jamais un filet (décision Lily 12/09) 2026-09-12 09:19:59 +02:00
state chore(state): do not version generated context-map (Lily, 08/09) 2026-09-08 14:38:46 +02:00
tests feat(cascade): le cloud est une décision, jamais un filet (décision Lily 12/09) 2026-09-12 09:19:59 +02:00
.gitignore chore(state): do not version generated context-map (Lily, 08/09) 2026-09-08 14:38:46 +02:00
config.json fix(search): renomme web_search → ivy_web_search (collision avec l'outil core) 2026-09-11 19:41:40 +02:00
context-map.review.md feat(context-map): étape 4 — ContextSelector (injection dynamique par sujet) 2026-08-16 18:48:31 +02:00
index.ts fix(search): renomme web_search → ivy_web_search (collision avec l'outil core) 2026-09-11 19:41:40 +02:00
keywords.json fix(search): renomme web_search → ivy_web_search (collision avec l'outil core) 2026-09-11 19:41:40 +02:00
Makefile chore(dataset,docs): cas I9/F3/secret-local, cible make compensates, textes périmés 2026-09-09 23:50:44 +02:00
openclaw.plugin.json fix(search): renomme web_search → ivy_web_search (collision avec l'outil core) 2026-09-11 19:41:40 +02:00
package-lock.json portage evaluate.mjs -> index.ts: before_model_resolve (route) avec classify() unique, categoryToRoute(), sticky v3 session (workCounter Map), pre-filtre, timeout 5s; tsconfig + @types/node; supprime ctx-check.mjs 2026-08-15 23:38:12 +02:00
package.json chore(dataset,docs): cas I9/F3/secret-local, cible make compensates, textes périmés 2026-09-09 23:50:44 +02:00
README.md fix(search): renomme web_search → ivy_web_search (collision avec l'outil core) 2026-09-11 19:41:40 +02:00
SPEC-v2.md fix(search): renomme web_search → ivy_web_search (collision avec l'outil core) 2026-09-11 19:41:40 +02:00
tsconfig.json feat(scaffold): milestone 1 — typed config, keywords, journal, no-drift install 2026-09-07 14:09:27 +02:00

ivy-router v2

OpenClaw plugin: sovereign-first router for Ivy. Local models handle everything they can; the cloud is reached only when justified, and only through a single fail-closed exit gate.

This branch (rewrite-v2) is a clean rewrite driven by SPEC-v2.md. Read the spec first — it is the source of truth, this README only describes what is built so far.

Status — milestone 9 of 10 (SPEC §17)

# Milestone State
1 Scaffold + config (typed loader, schema validation, keywords.json, no-drift install, logging + journal flag, capture killed) ✅ done
2 Typing (4 fields) + L0 socle + USER/MEMORY/TOOLS/AGENTS selectors + L2 assembly ✅ done
3 L4 exit gate (deterministic leak-net) ✅ done
4 L1 classifier + weight formula ✅ done
5 Cascade router ✅ done
6 RAG context-map: IDF scoring, freshness, tags lock ✅ done
7 delegate_work (nemotron + Claude CLI) + guardrails + persisted per-day budgets ✅ done
8 Image pipeline (§12) + web search & sources (§13) ✅ done
9 Dataset v2 (68 → 126 entries) + full non-regression run ✅ done
— Gateway hooks: L1→L4 wired to OpenClaw (src/gateway, 09/09 polish pass) ✅ done, never run against a live gateway yet
10 Test-LXC deploy + 7-day trial ⬜

The plugin loads its configuration, registers three tools (delegate_work, analyze_image, ivy_web_search) and the five gateway hooks that make it a router: model resolution, system-prompt assembly, the L4 gate on the final payload, the per-route tool allow-list and the gate on tool results. See "Gateway hooks" below. Everything is unit-tested against fakes; the live behaviour on the test LXC (milestone 10) is still to be observed.

Context engine (SPEC §8) — milestone 2

Every selectable unit of context carries four fields (SPEC §8.1): domain[], sensitivity (public | prive | jamais-sortant — exactly three levels, Lily 07/09), scope (socle | rappelable | archive) and valid (permanent | date | supersede). Sensitivity comes from code-level markers only; domain and valid come from config markers, because a mistake there costs relevance, not privacy.

L0 socle (src/context/socle.ts) is what goes out on every message: SOUL.md and IDENTITY.md in full (the persona is always complete — Lily 05/09), the hard operating rules from context.socle.operatingRules, and Lily's public mini-profile — the USER.md identity header only, never the ## Context body. It is public by construction: a header line that is not public is handed to the USER selector instead, where the route ceiling decides.

L2 selectors (src/context/selectors.ts) then add per-domain blocks under their own budgets: USER.md ## Context sub-sections, MEMORY.md base sections plus keyword matches, TOOLS.md by alias then keyword, AGENTS.md by keyword.

Three v1 bugs are fixed structurally here:

  • maxPersonaChars did nothing. Persona files were concatenated alphabetically and hard-cut at 12 000 chars, which truncated AGENTS.md while the whole USER.md — health section included — went out. The persona is now SOUL + IDENTITY only, in declared order, and the ceiling is checked and reported, never applied as a cut.
  • The selection loop break-ed on the first section that did not fit, silently dropping every candidate behind it — which is how USER.md's health context became all-or-nothing. It now continues, with a recorded reason per skip.
  • A section was included whole or not at all. Sections are now split per sensitivity and packed at line boundaries, so the public bullets of ### Core identity can reach a cloud model while its health bullet stays home.

Inspect an assembly without sending anything anywhere:

node src/cli.ts context "le tunnel ollama est tombé" local-action

It prints unit ids, sizes, sensitivities and skip reasons — never the content.

RAG context-map (SPEC §8.4) — milestone 6

A fifth selective source sits next to those four: the context-map, a deterministic haystack cut from the workspace files that have no selector of their own (context.contextMap.sources, e.g. skills/). Max 3 chunks, ≤ 4000 chars, no LLM anywhere in the path.

The v1 scoring was hits × (0.5 + importance), and the 05/09 review (context-map.review.md) showed why it had to go:

v1 v2 (src/context/rag.ts)
hits counted occurrences → the longest chunk won distinct keywords only, divided by √(chars / avgChars) — repeating a word earns nothing
every keyword weighed the same IDF: ln(1 + N / (1 + df)), so a corpus-wide word barely moves the score
one keyword was enough to be retrieved ≥ 2 distinct keywords (minDistinctKeywords) to be a candidate at all
everything that matched was ranked score floor (scoreFloor) — below it, a chunk is noise, not third place
importance/tags came from an LLM pass at 21 % agreement, and tags even gated the privacy filter tags are decorative: nothing in src/ reads them, and sensitivity is re-derived from code-level markers at use time

The map is versioned, timestamped, and hash-checked: every source file's sha256 is stored with it, and a rebuild happens when a byte changed or the format version moved — never because a clock moved. It lives in state/context-map.json, generated on each machine and gitignored (decision Lily 08/09: it holds verbatim workspace text), not hidden in the installed copy like v1's only copy was.

make map         # rebuild if a source changed
make map-check   # report staleness without writing (exit 1 if stale)

Assembling a payload detects staleness too and rebuilds in memory, so the models never see a stale map; writing the state file stays a maintenance action.

Gateway hooks (SPEC §5) — wired 09/09

src/gateway binds the layers above to OpenClaw. Every decision is a pure function of (deps, event, ctx) — tests/gateway.test.ts runs them against a scripted classifier and the fixture workspace — and index.ts only adds the failure posture (routing error → local brain, assembly error → minimal public prompt, gate error → block).

Hook What the router does there
before_model_resolve L1 classify (or image from the attachment), §7.2 decide, build the §7.1 plan, answer {providerOverride, modelOverride} for its first step. OpenClaw calls the hook again when a candidate fails: the next step is handed out, and once the plan is exhausted the router steps aside. Cron runs are never routed.
before_prompt_build The L2 assembly under the route's ceiling replaces the native system prompt (that replacement is the fix for v1's 72k chars per message). Nothing goes to prependContext: the user turn is Lily's message alone (avis 4B). Local routes get a final attribution anchor (COMPENSATES).
before_agent_run The L4 gate on the final payload, on the channel of the provider actually resolved (ctx.modelProviderId) — a native DeepSeek → Gemini fallback is gated with Gemini's public-only ceiling (v1 lesson C1). Unknown provider → block. Privacy-locked run on any cloud provider → block. Local payload over budget → block (the cloud compaction model must never see a local transcript). A secret bound for a third party → block, since this hook cannot redact.
before_tool_call routeCeilings.<route>.tools as an allow-list: ["*"] lifts it, [] forbids every tool. A route listing exec by name also filters the command against execCommands: one allowed script; no ; & | $ backtick < > or newline; no sudo/eval/sh -c; no env/elevated/host/node parameter.
tool_result_persist What a tool read mid-run crosses the gate before it re-enters the conversation (v1 lesson C2): refused → content-free notice, secret → redacted.
before_compaction Observation only: a compaction on a local-routed run is journaled.

The gateway must allow conversation hooks for this plugin (plugins.entries.router-dev.hooks.allowConversationAccess: true in openclaw.json), and the provider/model ids in config.json → models must exist in openclaw.json → models.providers.

Delegation & budgets (SPEC §10) — milestone 7

delegate_work(task, target?, model?, budgetHint?) hands one technical sub-task to a worker: nemotron locally (Ollama, default) or the Claude CLI (claude -p … --model opus, both bin and args configurable). What travels is the extracted sub-task, never the original message — "ce qui est relié au cerveau reste au cerveau, ce qui est relié au dev part au dev" (Lily 07/09).

What comes back is a tool result, not an answer (§9.4): the brain reads it and writes to Lily itself. Same for a refusal — it carries a sentence Ivy can say out loud, because a budget that degrades an answer silently is the failure mode §10.2 exists to prevent.

Trigger rules (src/delegation/triggers.ts) are config-driven and deterministic: explicit request, dev-heavy above directCloudThreshold, a local model signalling inability, or the keyword class — otherwise the local worker takes the dev categories and nothing else is delegated. They propose; they never permit. Permission is the ceiling check plus the L4 gate, in that order:

Guardrail Effect
route ceiling + jamais-sortant marker prive never reaches Claude (Q3, unanimous), jamais-sortant reaches nobody — not even the local worker
L4 gate on the claude-cli channel same gate, same grain, as every other outgoing action (§9.2) — secrets redacted, blocked ids refused, before any spend
depth + chain hash a worker cannot delegate for ever, nor re-delegate the task it was given
allowedModels opus by default, sonnet selectable, anything else refused rather than substituted
per-turn budget (RAM) one message's cascade cannot loop through 8 calls
per-day budgets (on disk) perDayClaudeCalls, perDayCloudCalls, perDayCostEur

The per-day counters live in state/budgets.local.json — gitignored, written temp-file-then-rename, re-read on every decision, reset on the UTC (or local) calendar day. That is the whole point: a counter guarding Lily's Claude Max quota must survive a restart, and one that lives in RAM gives the quota back on every Ctrl-C. An unreadable state file refuses the delegation instead of assuming zero.

make budgets     # today's counters against their ceilings

Image pipeline (SPEC §12) — milestone 8

analyze_image(path, message?, sensitivity?) picks one service deterministically from the message and the filename, then runs it behind the same gate as everything else:

Trigger Service Follow-up
"sauce", "source", "qui a dessiné", or a booru/pixiv-looking filename SauceNAO (upload, SAUCENAO_API_KEY) a Danbooru hit → Danbooru API for tags/characters
"quel épisode", "quel anime", a screenshot trace.moe (no key) —
anything else Gemini flash vision (GEMINI_API_KEY) —

Each service has its own declared L4 channel, so an image call is recognisable in the journal and an undeclared channel cannot emit at all (§9.2). What leaves as text is only the configured FR instruction — Lily's message stays with the brain (§10.1).

Privacy follows §12 exactly. The image's reference (path, filename) and the message around it are classified with the same frozen code-level markers the L4 gate reads: an image under a private folder, an NSFW path or the blocked Discord id is jamais-sortant and reaches no service — Ivy says so instead of sending it quietly. A merely prive image does go out, because decision 14 accepts that risk today — and it is journaled per image, with its sensitivity, which is what makes the risk an accepted one rather than an invisible one.

Quotas follow §10.3's three primitives: SauceNAO has a real ceiling, so its short (4 / 30 s) and long (100 / day) windows live in state/saucenao.local.json and survive a restart; Gemini vision is a cloud call and spends perDayCloudCalls; Danbooru and trace.moe are counted nowhere, because there is nothing there to protect. When SauceNAO cannot run — rate limited, no key, down — the pipeline degrades to a description and carries the reason, so the answer says "je n'ai pas pu chercher la source" instead of pretending it did.

Web search & sources (SPEC §13) — milestone 8

ivy_web_search(query, message?, category?, sensitivity?) runs the configured provider (search.provider: duckduckgo by default, gemini available) and comes back with the results and a ready-made sources block:

Sources :
Ollama — https://ollama.com/
Ollama — GitHub — https://github.com/ollama/ollama

The block is built in code rather than left to the brain's prompt, because decision 17 asks for it systematically — a model that forgets is a model that presents search results as its own knowledge. The tool result tells the brain to copy it verbatim at the end of the answer.

Two triggers (decision 16): Lily asks explicitly (cherche, vérifie, actualité…), or Ivy is unsure about a fact. One lock beats both: no search on private or personal topics. That lock is the existing privacy mechanism — classifyText over the frozen code-level markers, widened (never narrowed) by keywords.json — applied to the query and to the message it came from, so a neutral-looking query extracted from a health message is still refused. The refusal is spoken, not silent.

The query then crosses the L4 gate on the web-search channel like any other outgoing payload, so a secret pasted into a message is redacted before it reaches a search engine. DuckDuckGo counts nothing; Gemini-as-search spends perDayCloudCalls (§10.3 — its ceiling is a quota, and a stall is silent).

The name is ivy_web_search, not web_search (11/09): OpenClaw's core web_search wins a name conflict and the gateway skips the plugin's tool, so under the core name none of the above ever ran — the query went to the core provider with no privacy lock, no L4 and no sources block.

Layout

ivy-router/
├── SPEC-v2.md              # source of truth
├── index.ts                # OpenClaw plugin entry (wiring only)
├── config.json             # ALL tunables — schema-validated at load
├── keywords.json           # SPEC §11 intent tables + additive privacy keywords
├── Makefile                # install (symlink, no drift) / check / test
├── src/
│   ├── paths.ts            # repo-root & ~ resolution (never CWD-relative)
│   ├── runtime.ts          # composition root: config + keywords + journal
│   ├── cli.ts              # `make check` + `context` inspector
│   ├── config/             # schema validator, config schema, loader, types
│   ├── context/            # SPEC §8: typing, L0 socle, selectors, L2 assembly
│   ├── classifier/         # SPEC §7.2 L1: smollm2 call + constrained parsing
│   ├── routing/            # SPEC §7.2: weight formula, sticky session counter
│   ├── cascade/            # SPEC §7.1: route plan, fallback order, degradation
│   ├── budgets/            # SPEC §10.2: per-day counters (disk) + per-turn (RAM)
│   ├── delegation/         # SPEC §10: delegate_work, triggers, guardrails
│   ├── media/              # SPEC §12: image pipeline, service selection, rate limit
│   ├── search/             # SPEC §13: ivy_web_search, privacy lock, sources block
│   ├── gateway/            # SPEC §5: the OpenClaw hooks (resolve, prompt, L4 gate, tools)
│   ├── egress/             # the only place src/ faces the outside (L4 + transports)
│   ├── keywords/           # keywords.json loader + code-level privacy floor
│   ├── dataset/            # SPEC §16.1: dataset v2 loader + non-regression replay
│   └── logging/journal.ts  # opt-in, rotating, content-free journal
├── state/                  # router state (generated, gitignored: context-map + local counters)
├── tests/                  # node:test — `make test`
├── legacy/                 # v1 source, reference only (see legacy/README.md)
├── scripts/                # dev tools (context-map chunking, evaluation, dataset runner)
├── data/dataset.jsonl      # v1 dataset, untouched: 68 × (message → categorie)
└── data/dataset-v2.jsonl   # dataset v2: those 68 + privacy/cascade/keyword cases

Config (SPEC §14)

config.json holds every tunable: model topology, weight formula, thresholds, per-domain context budgets, route ceilings, delegation guardrails, journal settings. It is validated against a schema at load, and the loader throws rather than falling back to defaults — v1 ran for weeks on a config nobody knew was wrong. Unknown keys are errors too, so a typo cannot be silently ignored; _-prefixed keys are comments.

Beyond the schema, the loader enforces invariants that encode locked decisions: routing weights must sum to 1, localLock >= directCloud, every category has a route, private routes local-only, no route may declare jamais-sortant as its ceiling (that content enters no prompt at all), and a cloud text route may not exceed public — the cloud only ever receives public content (Lily 07/09). Milestone 8 adds its own: a URL template must keep its placeholder ({id}, {model}), the SauceNAO rate-limit state must be named *.local.json so git never tracks it, and its short window must actually be shorter than the long one. The gateway layer adds one more: the local providers, the DeepSeek provider and the other cloud providers must be three disjoint sets, because the L4 channel of a run is derived from the provider id alone.

routeCeilings.<route>.tools is enforced at execution time by before_tool_call: ["*"] means no restriction, [] means no tool at all — an empty list is never "anything goes". The Notion/Obsidian capabilities are ~/.openclaw/bin scripts run through exec, so both local routes list exec and restrict it with routeCeilings.<route>.execCommands (11/09, C5 — mapping in _comment, config.json). The shipped defaults are provisional and need Lily's call.

Secrets are never in this repo — only ~/.openclaw/.env (GEMINI_API_KEY, SAUCENAO_API_KEY, DEEPSEEK_API_KEY…). The plugin reads the environment, never the config.

Privacy keywords are additive

keywords.json carries the SPEC §11 intent tables and a privacyKeywords list. That list is unioned with the code-level floor in src/keywords/privacy.ts, so editing or emptying the JSON can only ever widen the local lock, never narrow it.

Logging (SPEC §15)

The v1 prompt capture is gone. It wrote every system prompt, message and history entry to a single JSONL file that reached 463 MB.

Its replacement is opt-in and records decisions, not content: routing decisions, L4 refusals/redactions by rule id, delegation events, external API calls (image services and web searches: service, file/query hash, result count, and the reason when the call did not happen). Content-bearing keys are stripped recursively before writing, and a journal failure never propagates into a run.

Turn it on with --journal, IVY_ROUTER_JOURNAL=1, or journal.enabled in config.json (CLI flag > env > config). Files rotate by size — 10 MB × 3 by default — in ~/.openclaw/router-journal/.

To delete the v1 capture:

make purge-capture              # shows what would go
make purge-capture CONFIRM=1    # deletes it

Note the capture has a producer outside this repo: the separate prompt-capture plugin. Disable it in ~/.openclaw/openclaw.json (plugins.entries.prompt-capture.enabled = false) or the file grows back.

No-drift install (SPEC §14)

In v1 the installed plugin was a second git clone that had diverged from the repo: four config.json.bak-* files, three stray test scripts, and the only copy of context-map.json. make install replaces that with a symlink, so git checkout is the only way the installed plugin can change.

make install     # ~/.openclaw/plugins/ivy-router -> this repo
make status      # what is installed, and whether it drifted
make uninstall   # removes the symlink (never a real directory)

make install refuses to overwrite a real directory — back the v1 copy up first (mv ~/.openclaw/plugins/ivy-router{,.v1-backup}).

Development

make check       # validate config.json + keywords.json, print a summary
make test        # unit tests (node:test, native TypeScript)
make dataset     # replay the non-regression dataset, entry by entry
make typecheck   # tsc --noEmit

Per SPEC §4.1: development happens on the Ivy LXC (~/Labs); real runtime tests happen on the test LXC.

Non-regression dataset (SPEC §16.1)

data/dataset-v2.jsonl extends the v1 68-entry dataset with the three axes §18.6 asked for — privacy cases that must stay local, cascade cases that name their fallback target, keyword→tool cases — and each entry states what the deterministic route-resolution core must produce: §7.2 route decision, §7.1 cascade plan and fallback, §9.2 L4 verdict per channel, §11 intents, §13 search lock.

The classifier is an input, not a call: an entry carries its categorie and confidence, so the replay is offline, deterministic and part of make test. The v1 file stays byte-identical, and the runner refuses a v2 set that lost or relabelled any of its 68 entries.

npm run dataset                  # or: make dataset
npm run dataset -- --kind=privacy --verbose
npm run dataset -- --live        # opt-in: 5 runs/entry against the real classifier

--live is the §16.1 evaluation protocol (temperature 0, 5 runs, read the distribution and not the last draw). It needs Ollama serving models.classifier on the machine it runs on — the Ivy LXC — and it prints without gating: an unstable 500M model is an open bug to look at, never a red deterministic suite.

Entries whose expectation records what the code does today where SPEC-v2 asks for something else carry a divergence id, documented in src/dataset/divergences.ts and printed on every run. A pinned divergence is a decision waiting for Lily, not a passing test.

The list is empty since 08/09: the four divergences M9 pinned (the §11 route lock, the unreachable directCloud score branch, bare nsfw at the gate, the vision channels' declared ceiling) were decided and fixed in the code, and the entries that pinned them now assert the corrected behaviour.

  • Ivy-AI/ivy-persona — identity + memory
  • Ivy-AI/ivy-infra — scripts + config (cron, skills, bin, systemd)
  • Ivy-AI/ivy-sessions — raw conversations (.jsonl)
  • Ivy-AI/ivy-hub — Android Flutter app
  • Ivy-AI/ivy-rag — RAG on notes (LanceDB)