What it is
INTERPLANE belongs to neither the model side nor the runtime side. Runtimes keep their own tools, policy, approvals, memory and execution. INTERPLANE only standardizes what crosses between a model and a runtime, and makes it measurable.
- Interplane Probe
- Measured capability qualification of an OpenAI-compatible endpoint.
- Lenshift
- Model dialect parsers and result renderers (
openaiandqwen35today;ajaxis reserved). - CrossAxis
- Explicit, versioned capability mapping, plus a deterministic domain selector that records a before and after receipt.
- Crossveil
- The authority-boundary contract: two runtime callbacks, a lifecycle with fail-closed invariants, and trust classes.
- RelayLine
- The event vocabulary for multi-round exchanges (a contract only so far).
- Vectorveil
- The envelope and canonical JSON (envelope only so far; no transport framework).
- Bench
- The benchmark harness: a paired runner and analyzer that compare a full tool catalog against a CrossAxis selection under a pre-registered protocol.
Security in one paragraph: the model side is untrusted. Lenshift never executes. CrossAxis never grants. A valid envelope means "structurally understandable", never "authorized". The runtime makes every authorization decision through its own policy engine, and a denial never becomes an authorization by retry.
Status
Every number below comes from the repository's own reports. The repository's STATUS.md is the authority on what is implemented, tested or only specified.
0.1, Reality layer: goal met
- Normative spec: 10 JSON Schema (2020-12) documents plus the written specification.
- Rust and Python implementations produce byte-identical conformance verdicts on all cases in the 0.1 report (CI job
cross-language). - Goal receipt, same Qwen3.5 task over a 71-tool catalog: CrossAxis selected 9 tools. First-turn prompt tokens fell from 15,846 to 1,674 (89.4% fewer). Both runs reached the correct answer, and a second independent run reproduced the selection and the token reduction.
- Probe on Qwen3.5-9B with Ollama 0.34.0: 13 pass, 1 unsupported (
tools.text_qwen35), with identical verdicts from the Rust and Python probes. - Stated as unproven in the report: any claim about Ajax (no official artifact exists), and a task-success benchmark (only one task was measured at 0.1).
0.2, Capability benchmark: pre-registered gate FAILED
The full 0.2 report is still being written; this section shows only results already published in the repository.
The protocol was pre-registered before any model run. The qualification run used 37 tasks, each run once with the full 71-tool catalog (A) and once with CrossAxis Select (B), on qwen3.5:9b. Overall result: FAIL.
| Criterion | Measured | Threshold | Result |
|---|---|---|---|
| T: median tool-schema token reduction | 0.877 | ≥ 0.7 | pass |
| S: success not worse | success 0.676 (A) vs 0.730 (B); CI lower bound -0.085 | lower bound ≥ -0.10 | pass |
| O1: required tools exposed (final) | 0.865 | ≥ 0.95 | fail |
| O2: recovered within 1 expansion | 0 of 5 | ≥ 0.9 | fail |
The O1 and O2 failures come from the five tasks built to need an expansion after round one. In condition B the model made one call to the discovery tool across all 37 runs, and the required expansion did not occur on any of those five tasks. Those five tasks also failed under the full catalog. The tool-cost reduction held and success was not measurably worse, but the claim as pre-registered is not supported. The result applies only to this model, backend and machine.
Backends qualified for Qwen3.5-9B
- Ollama 0.34.0: 13 pass, 1 unsupported (Ollama's OpenAI path consumes the text-form tool call).
- llama.cpp b11398: all 14 probe checks pass.
- SGLang 0.5.20: all 14 probe checks pass in both the Python probe and the Rust probe. The Rust probe was rerun after the two probes were aligned (interplane main commit
e34160f); the earlier Rust run is kept in the repository as a pre-alignment record.
Roadmap
| Version | Layer | Goal |
|---|---|---|
| 0.2 | Capability | Give small local models the minimum tools required without making them less capable. |
| 0.3 | Trust | Preserve the line between content, model intent and runtime authority. |
| 0.4 | Execution | Multi-round execution without losing the boundary. |
| 0.5 | Qualification | Measured, reproducible model and backend claims. |