Files
Labvoice/planning/mcp-server-plan.md
T
2026-08-30 20:44:00 +02:00

442 lines
23 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# LabVoice MCP Server — Implementation Plan
Status: **draft for review** (nothing below is built yet; `src/` is empty and
`vendor/elabapi-python` holds the vendored official `elabapi_python` client)
## 1. Goal
An MCP server exposing the eLabFTW REST API v2 as a small set of **aggregate tools**
designed for small/low-power models and voice-driven wetlab use.
Design decisions already made:
1. **Consumables live in the protocol template** — machine-readable JSON at the end of
the step text, hidden from the client UI in an HTML comment.
2. **Direct execution** — no preview/confirmation round. The point is hands-free
voice operation.
3. **Portable** — templates reference stable *semantic resource keys*, never eLabFTW
numeric IDs. The server installs on top of any existing eLabFTW instance with only
template edits (plus optional resource markers).
4. **Reads are dual-purpose** — the same read tools are exposed both as MCP tools and
over a plain read-only REST API. REST lets the mobile app render state without
spending model tokens; MCP lets the LLM fetch detail on demand. The edge LLM
receives only a **minimal handoff context** from the app (experiment id, user id,
current step id, device id, location id — on-device compute means a tight token
budget), so it pulls anything beyond the ids via the read tools.
5. **Templates are first-class** — new experiments are created from a template, chosen
interactively via a template-listing/search tool (short descriptions, semantic
ranking).
## 2. Architecture
```text
mobile app ──REST: reads for the UI (no LLM)──┐
│ │
│ voice + minimal handoff │
│ context (ids only) │
▼ │
small LLM ──MCP: reads on demand, aggregates──┤
(stdio / streamable HTTP) │
▼
┌──────────────────────────────┐
│ mcp_server │
│ ├── tools.py │ small tool surface, compact results
│ ├── rest.py │ read-only REST mirror of the reads
│ ├── templates.py │ template search index + creation
│ ├── workflows/ │ aggregate operations (the point)
│ │ ├── protocol.py │ step completion saga
│ │ └── inventory.py │ allocation, unit handling
│ ├── resolve.py │ resource_key → item id mapping
│ ├── annotate.py │ annotation parse/validate
│ ├── elabftw.py │ async adapter over vendored elabapi_python
│ └── journal.py │ SQLite operation journal (sagas)
└──────────────────────────────┘
│ HTTPS, Authorization: <api key>
▼
eLabFTW REST API v2
(client: vendor/elabapi-python)
```
Layers, top to bottom: **MCP tools / REST routes → workflow services →
elabapi_python adapter → eLabFTW**. The read tools are single implementations behind
two transports: `tools.py` (MCP, for the LLM) and `rest.py` (REST, for the app UI).
No generic `PATCH/POST/DELETE` passthrough tools — aggregate tools only, so a small
model cannot pick an unsafe raw mutation.
## 3. Module layout (`src/mcp_server/`)
```text
src/mcp_server/
__init__.py
__main__.py # entry point: stdio (default) or streamable-http
config.py # env-driven settings (pydantic-settings)
models.py # tool input/output schemas, compact result types
errors.py # typed errors → voice-friendly messages
elabftw.py # async adapter over the vendored elabapi_python client:
# Configuration/ApiClient setup, error normalization,
# sync calls wrapped with asyncio.to_thread
schemas.py # pydantic models for the eLabFTW entities we use
annotate.py # labvoice:v1 annotation parser + validator
resolve.py # resource key resolution (mapping store + markers)
allocate.py # container selection, unit whitelist/conversion
templates.py # experiment templates: description extraction, semantic
# ranking (embedding cache in SQLite), creation flow
journal.py # SQLite saga journal (idempotency + compensation)
workflows/
__init__.py
protocol.py # complete_next/complete/observation workflows
inventory.py # adjust_inventory, stock reads
setup.py # scan_instance, map_resource, validate_templates
tools.py # MCP tool registration (one function per tool)
rest.py # FastAPI: read-only REST mirror of the read tools
# (state rendering for the app UI; bearer-token protected)
server.py # FastMCP assembly, lifespan; mounts rest.py in http mode
vendor/elabapi-python/ # official generated client (package elabapi_python,
# sync urllib3) — vendored, not edited; imported as-is
```
Dependencies to add: `mcp` (official Python SDK), `pydantic-settings`,
`aiosqlite` (journal), `fastapi` + `uvicorn` (REST fast path), `fastembed`
(optional — semantic template ranking, unset ⇒ lexical fallback). The REST client
itself comes from the vendored `elabapi_python` package
(`elabapi_python.Configuration` with `host = ELABFTW_URL + /api/v2`, api-key auth
header); because it is synchronous, the adapter runs calls via `asyncio.to_thread`.
Python 3.12 (already pinned).
## 4. Protocol annotation format (portable)
Appended at the end of a template step body:
```html
<!-- labvoice:v1
{
"consumables": [
{"resource_key": "ethanol_absolute", "quantity": 2.0, "unit": "mL",
"allocation": "fifo"}
]
}
-->
```
Field rules:
| field | required | notes |
|-----------------|----------|----------------------------------------------------|
| `resource_key` | yes | slug; stable across instances — never an item id |
| `quantity` | yes | positive number consumed per execution |
| `unit` | yes | must be in the unit whitelist |
| `allocation` | no | `fifo` (default), `nearest_expiry`, `specific` |
| `container_id` | no | only with `allocation: "specific"`; local hint, ignored if it does not match the resolved resource |
| `optional` | no | `true` → skip (with a warning) if stock is missing |
Parser rules (`annotate.py`):
- Recognize only `<!-- labvoice:v1 ... -->`; take the **last** valid block.
- Malformed JSON or schema violations ⇒ `annotation_error`, surfaced by
`validate_protocol_template` and blocking step completion (never silently ignored).
- `quantity` may also be `null` with `"prompt_quantity": true` for steps where the
used amount varies — completion then **requires** the caller to supply quantities
(voice: “how much did you use?”), otherwise it fails with a clarification request.
- Visible step text is never modified by the server.
Unit whitelist (extensible in config): `μL, mL, L, mg, g, kg, μg`, `ea` (each).
Conversions only within the same dimension and only tested pairs (e.g. `mL↔L`,
`mg↔g↔kg↔μg`); anything else ⇒ clarification, never a guess.
## 5. Resource resolution (`resolve.py`)
`resource_key` → eLabFTW `items` id, resolved in this order:
1. **Local mapping store** (SQLite table): `resource_key → item_id`, written during
setup. Authoritative at execution time.
2. **Resource marker** — hidden comment in the resource body:
`<!-- labvoice:resource-key=ethanol_absolute -->`, auto-discovered by scanning.
3. **Configured matchers** at setup time only (CAS extra field, `custom_id`).
4. **Exact title match** — setup-time suggestion only, requires admin approval;
never applied silently during execution.
Ambiguous or missing mapping ⇒ execution stops with a voice-friendly clarification
listing candidate resources. No fuzzy matching at runtime, ever.
### Installation on an existing instance (setup workflow)
- `scan_instance` — walk `GET /items` (paginated), inventory containers, detect
markers, propose matches from title/CAS/custom_id.
- `map_resource` — bind `resource_key → item_id` (idempotent, upsert).
- `export_mapping` / `import_mapping` — JSON mapping file for porting between
instances; import proposes but still requires approval of ambiguous matches.
Mapping file example:
```json
{
"ethanol_absolute": {"cas": "64-17-5", "title": "Absolute Ethanol", "unit": "mL"}
}
```
## 6. MCP tool surface
Ten tools total. The read tools are dual-purpose: the mobile app calls them via the
REST mirror (§8) to render state with zero model tokens, while the edge LLM — which
receives only a minimal handoff context of ids from the app — calls the same reads
over MCP when it needs detail (e.g. `get_next_protocol_step` for the full step body
and stock preview).
Reads (safe, any API key):
| tool | purpose |
|-------------------------|--------------------------------------------------------|
| `find_experiments` | search by title/text/tag/custom id, compact results |
| `get_experiment_context`| metadata + unfinished steps + linked resources + recent comments, all compact |
| `get_next_protocol_step`| next unfinished step: text, parsed consumables, stock preview |
| `list_experiment_templates` | list/search templates: id, title, short description, tags (semantic ranking, §9) |
| `validate_protocol_template` | check annotations + mappings of a template/experiment |
Aggregates (mutating):
| tool | purpose |
|-----------------------------|-------------------------------------------------------------------------|
| `complete_next_protocol_step` | the headline tool — see §7 |
| `complete_protocol_step` | same, with explicit step id (for “redo step 3” voice commands) |
| `record_protocol_observation` | add a comment (+ optional step body edit) without completing anything |
| `create_experiment_from_template` | start a new experiment from a template (interactive flow, §9) |
| `adjust_inventory` | restock / correct a container, voice: “add 500 mL to …” |
Setup (admin, mutating only the local mapping store):
| tool | purpose |
|---------------|---------------------------------------------|
| `scan_instance`, `map_resource`, `export_mapping`, `import_mapping` | §5 onboarding |
Result shapes are deliberately small: every tool returns compact JSON (ids, titles,
quantities, statuses) — never a raw eLabFTW entity dump. Tool descriptions are one or
two short sentences (small-model friendly). Mutating tools return exactly what
changed, for text-to-speech readback:
```json
{
"ok": true,
"experiment_id": 123,
"step": {"id": 9, "body": "Add ethanol", "finished": true},
"consumed": [{"resource_key": "ethanol_absolute", "container_id": 12,
"amount": "2.0 mL", "remaining": "48.0 mL"}],
"comment_id": 77,
"next_step": {"id": 10, "body": "Incubate 30 min"}
}
```
## 7. `complete_next_protocol_step` workflow (direct execution)
Input: `{experiment_id, comment?, quantities?}` — nothing else.
1. `GET /experiments/{id}` + steps; select lowest-`ordering` unfinished step.
None left ⇒ explicit `protocol_complete` result.
2. Parse annotation (§4). No annotation ⇒ complete step + comment, skip inventory.
3. Resolve every `resource_key` (§5). Unresolvable ⇒ clarification result, no mutation.
4. `GET /{entity_type}/{id}/containers` per resource; filter by unit compatibility;
apply allocation policy (`fifo` = lowest container id with stock,
`nearest_expiry` needs an expiry extra field, configured).
5. Validate stock: total available ≥ required (unit-converted). Insufficient ⇒
clarification listing what is short; `optional` consumables are skipped with a warning.
6. Execute as a journaled saga (§10), in this order:
1. decrement each container (`PATCH .../containers/{subid}` `qty_stored`),
2. finish the step (`PATCH .../steps/{subid}` `{"action":"finish"}`),
3. post the comment (`POST .../comments`).
7. Re-read the step to verify; return compact confirmation + next step.
The model makes **one tool call**; all sequencing is server-side.
## 8. Read-only REST API (app fast path)
The read tools are single implementations behind two transports: `rest.py` exposes
them as REST endpoints so the mobile app renders state without spending LLM tokens,
while `tools.py` keeps them available over MCP for the LLM and standalone clients.
Mutations are **not** exposed over REST — they exist only as MCP tools, so every
mutation is journaled (§10).
- Auth: static bearer token (`LABVOICE_REST_TOKEN`) — the app is a trusted client;
per-device tokens are an open question (§15).
- Same compact result shapes and typed errors as the MCP tools
(`{"error": "clarification", "message": ...}`).
| endpoint | purpose |
|---------------------------------------|---------------------------------------------------------------|
| `GET /api/experiments?q=&limit=` | `find_experiments` |
| `GET /api/experiments/{id}/state` | full app state: title, unfinished steps, linked resources, next step + parsed consumables + stock preview |
| `GET /api/experiments/{id}/next-step` | just the next step (the hot path for voice) |
| `GET /api/templates?q=&limit=` | `list_experiment_templates` (§9) |
The `state` response is what the app renders (title, steps, stock, etc.). What the
app forwards to the LLM is deliberately **minimal** — a handoff context of ids only,
since LLM compute is on-device at the edge and the token budget is tight:
```json
{"experiment_id": 123, "user_id": 2, "step_id": 9,
"device_id": "bench-7", "location_id": "lab-2"}
```
When the model needs more than the ids — full step body, parsed consumables, stock,
comments — it calls the corresponding MCP read tool (same services that back the
endpoints above). REST serves the UI; MCP serves the model; one implementation of
each read.
## 9. Experiment templates: semantic search & creation
Voice flow ("start a new experiment for the PCR cleanup"):
1. `list_experiment_templates(query?, limit?)` — templates as
`{id, title, short_description, tags}`, ranked by semantic similarity when a
query is given (best match first).
2. The LLM reads the short descriptions back; the user picks one interactively.
3. `create_experiment_from_template(template_id, title?)` — `POST /experiments`
with the template id; returns `{experiment_id, title, first_step}`. eLabFTW
copies the template steps, so `labvoice:v1` annotations come along and the
step-completion workflow (§7) applies immediately.
Short description: first paragraph of the template description, truncated
(`LABVOICE_TEMPLATE_DESC_LIMIT`, default 200 chars). It doubles as readback text
for the LLM and as part of the search corpus.
Semantic ranking (`templates.py`):
- Corpus per template: title + tags + full description.
- Preferred backend: local tiny embedding model (`fastembed`, ONNX, CPU, ~30 MB,
`LABVOICE_EMBED_MODEL`); template vectors computed at scan time (or lazily on
first use) and cached in the SQLite database as `template_id → vector`,
refreshed when templates change.
- Fallback (no model configured): lexical scoring — weighted token overlap,
title > tags > description. Same interface, just dumber ranking.
- Only templates visible to the current API key are ever returned; no fuzzy
matching on ids.
Creation notes:
- A single `POST /experiments` — no multi-step saga; still journaled for audit
(§10).
- `title` optional — eLabFTW applies the template's default title format when
omitted.
- A created-but-unwanted experiment is archived via eLabFTW itself; this server
never deletes.
## 10. Saga journal & failure handling (`journal.py`)
eLabFTW has no cross-entity transactions, so every mutating workflow runs as a
journaled saga in SQLite (`LABVOICE_DB_PATH`, default `~/.labvoice/journal.sqlite`):
- Each execution gets an `operation_id`; journal rows record planned actions, their
status, and eLabFTW responses.
- Sub-actions are idempotent (re-read before write; finishing an already-finished
step is a no-op; decrement uses read-modify-write with a re-read check).
- On failure of step 6.2 or 6.3: **compensate** — restore decremented quantities
(PATCH back), then:
- success ⇒ return `reverted` result explaining what happened;
- compensation fails ⇒ mark `partial_failure` in the journal **and** post an audit
comment on the experiment describing exactly what is inconsistent; result tells
the user which container to check.
- Journal is also the audit log (who/when/what) and powers a future reconciliation tool.
## 11. Configuration & deployment
Env vars (all via `config.py`):
```text
ELABFTW_URL # https://eln.example.org (adapter sets host to this + /api/v2)
ELABFTW_API_KEY # read-only works for read tools; writes need can_write
ELABFTW_TIMEOUT=10 # per-request timeout seconds
ELABFTW_RETRIES=2
LABVOICE_DB_PATH=~/.labvoice/journal.sqlite
LABVOICE_REST_TOKEN=... # bearer token for the read-only REST API (§8)
LABVOICE_TEMPLATE_DESC_LIMIT=200
LABVOICE_EMBED_MODEL=... # optional; unset ⇒ lexical template ranking
LABVOICE_UNIT_WHITELIST=... # optional override
LABVOICE_EXPIRY_FIELD=... # extra-field name for nearest_expiry allocation
```
- Transports: **stdio** (default; local/phone use) and **streamable-http**
(`--transport http`, for a shared Raspberry-Pi-class deployment). In http mode a
single uvicorn process serves both the MCP app and the read-only REST API
(`rest.py`) — one deployment for app + LLM. `stateless_http` mode; no session
affinity needed.
- Health check = `GET /info` through the client (auth + reachability in one call).
- Packaging: uv project, console script `labvoice`; single Dockerfile (optional).
## 12. eLabFTW endpoints used
```text
GET /info
GET /experiments (search: q, tags, limit/offset)
GET /experiments/{id}
POST /experiments {"template": <id>, "title"?} (§9 creation)
GET /{entity_type}/{id}/steps
PATCH /{entity_type}/{id}/steps/{subid} {"action": "finish"}
GET /{entity_type}/{id}/comments
POST /{entity_type}/{id}/comments
GET /items (scan/search)
GET /items/{id}
GET /{entity_type}/{id}/containers
PATCH /{entity_type}/{id}/containers/{subid} {"qty_stored": ...}
GET /storage_units?hierarchy=true (location names for readback)
GET /experiments_templates (list/search templates, §9)
GET /experiments_templates/{id}
```
`entity_type` ∈ {`experiments`, `items`} only — templates are read, never mutated, by
this server (template editing happens in eLabFTW itself). The adapter (`elabftw.py`)
normalizes all client errors into typed errors (`auth_error`, `permission_error`,
`not_found`, `api_error`) with voice-friendly messages; never leaks the API key into
results or logs.
## 13. Testing & verification
- **Unit** (pytest, mocked `elabapi_python` API instances): annotation parser (valid/malformed/
last-block/prompt_quantity), resolver precedence + ambiguity, allocator (fifo,
splitting across containers, unit conversion, insufficient stock), journal
(idempotency, compensation, partial_failure), error normalization, template
ranking (golden queries for semantic + lexical fallback, description truncation),
creation from template (steps + annotations copied).
- **REST API**: FastAPI TestClient — bearer auth enforced, `state`/`next-step`
responses (snapshot tests).
- **MCP-level**: tool schema lint (small input schemas), `tools/list` snapshot test,
golden results for the headline workflow.
- **Integration** (optional, manual): disposable eLabFTW docker instance; script
creates experiment + template with annotations, maps resources, runs the saga,
asserts final state via the API.
- **Safety checks as tests**: read-only key must fail cleanly on every mutating tool;
immutable step → explicit error; double execution of same operation_id is idempotent.
## 14. Build order (PR-sized milestones)
1. `config.py`, `errors.py`, `models.py`, elabftw adapter over `elabapi_python` + schemas (read paths only)
2. `annotate.py` parser + `validate_protocol_template` tool
3. `resolve.py` mapping store + `scan_instance` / `map_resource` tools
4. read tools: `find_experiments`, `get_experiment_context`, `get_next_protocol_step`
5. `rest.py`: read-only REST mirror over the same services (bearer auth)
6. `journal.py` saga engine
7. `workflows/protocol.py`: `complete_protocol_step` → `complete_next_protocol_step`
→ `record_protocol_observation`
8. `templates.py`: `list_experiment_templates` (search index) +
`create_experiment_from_template`
9. `adjust_inventory`
10. MCP server assembly (stdio + http, REST mount), entry point, README, Dockerfile
11. integration script + docs (annotation authoring guide for template editors)
Milestones 1–5 are read-only and independently testable; the first mutating code
lands in milestone 7 on top of the journal (6).
## 15. Open questions (need your input before/at build time)
1. **Step selection** — strictly lowest `ordering` among unfinished steps, or should
`complete_protocol_step` also accept a step *number* spoken by the user (“done
with step three”)? (Plan assumes yes: accept id or 1-based position.)
2. **`nearest_expiry`** — is an expiry extra field available in your resources, or is
`fifo` + `specific` enough for v1? (Plan: ship `fifo`/`specific` first.)
3. **Journal location on phone/mobile** — default `~/.labvoice/journal.sqlite` okay,
or should the SQLite file live next to the config for easy backup?
4. **MCP SDK line** — plan targets the current stable `mcp` SDK (v2 line). Pin major
version at build time.
5. **Template search backend** — ship lexical ranking first and add `fastembed`
embeddings only if ranking disappoints, or embed from day one? (Plan: lexical
first, same interface for both.)
6. **REST auth** — one shared bearer token per deployment enough, or per-device
tokens for app installs?