# LabVoice MCP Server — Implementation Plan Status: **draft for review** (nothing below is built yet; `src/` is empty and `vendor/elabapi-python` holds the vendored official `elabapi_python` client) ## 1. Goal An MCP server exposing the eLabFTW REST API v2 as a small set of **aggregate tools** designed for small/low-power models and voice-driven wetlab use. Design decisions already made: 1. **Consumables live in the protocol template** — machine-readable JSON at the end of the step text, hidden from the client UI in an HTML comment. 2. **Direct execution** — no preview/confirmation round. The point is hands-free voice operation. 3. **Portable** — templates reference stable *semantic resource keys*, never eLabFTW numeric IDs. The server installs on top of any existing eLabFTW instance with only template edits (plus optional resource markers). 4. **Reads are dual-purpose** — the same read tools are exposed both as MCP tools and over a plain read-only REST API. REST lets the mobile app render state without spending model tokens; MCP lets the LLM fetch detail on demand. The edge LLM receives only a **minimal handoff context** from the app (experiment id, user id, current step id, device id, location id — on-device compute means a tight token budget), so it pulls anything beyond the ids via the read tools. 5. **Templates are first-class** — new experiments are created from a template, chosen interactively via a template-listing/search tool (short descriptions, semantic ranking). ## 2. Architecture ```text mobile app ──REST: reads for the UI (no LLM)──┐ │ │ │ voice + minimal handoff │ │ context (ids only) │ ▼ │ small LLM ──MCP: reads on demand, aggregates──┤ (stdio / streamable HTTP) │ ▼ ┌──────────────────────────────┐ │ mcp_server │ │ ├── tools.py │ small tool surface, compact results │ ├── rest.py │ read-only REST mirror of the reads │ ├── templates.py │ template search index + creation │ ├── workflows/ │ aggregate operations (the point) │ │ ├── protocol.py │ step completion saga │ │ └── inventory.py │ allocation, unit handling │ ├── resolve.py │ resource_key → item id mapping │ ├── annotate.py │ annotation parse/validate │ ├── elabftw.py │ async adapter over vendored elabapi_python │ └── journal.py │ SQLite operation journal (sagas) └──────────────────────────────┘ │ HTTPS, Authorization: ▼ eLabFTW REST API v2 (client: vendor/elabapi-python) ``` Layers, top to bottom: **MCP tools / REST routes → workflow services → elabapi_python adapter → eLabFTW**. The read tools are single implementations behind two transports: `tools.py` (MCP, for the LLM) and `rest.py` (REST, for the app UI). No generic `PATCH/POST/DELETE` passthrough tools — aggregate tools only, so a small model cannot pick an unsafe raw mutation. ## 3. Module layout (`src/mcp_server/`) ```text src/mcp_server/ __init__.py __main__.py # entry point: stdio (default) or streamable-http config.py # env-driven settings (pydantic-settings) models.py # tool input/output schemas, compact result types errors.py # typed errors → voice-friendly messages elabftw.py # async adapter over the vendored elabapi_python client: # Configuration/ApiClient setup, error normalization, # sync calls wrapped with asyncio.to_thread schemas.py # pydantic models for the eLabFTW entities we use annotate.py # labvoice:v1 annotation parser + validator resolve.py # resource key resolution (mapping store + markers) allocate.py # container selection, unit whitelist/conversion templates.py # experiment templates: description extraction, semantic # ranking (embedding cache in SQLite), creation flow journal.py # SQLite saga journal (idempotency + compensation) workflows/ __init__.py protocol.py # complete_next/complete/observation workflows inventory.py # adjust_inventory, stock reads setup.py # scan_instance, map_resource, validate_templates tools.py # MCP tool registration (one function per tool) rest.py # FastAPI: read-only REST mirror of the read tools # (state rendering for the app UI; bearer-token protected) server.py # FastMCP assembly, lifespan; mounts rest.py in http mode vendor/elabapi-python/ # official generated client (package elabapi_python, # sync urllib3) — vendored, not edited; imported as-is ``` Dependencies to add: `mcp` (official Python SDK), `pydantic-settings`, `aiosqlite` (journal), `fastapi` + `uvicorn` (REST fast path), `fastembed` (optional — semantic template ranking, unset ⇒ lexical fallback). The REST client itself comes from the vendored `elabapi_python` package (`elabapi_python.Configuration` with `host = ELABFTW_URL + /api/v2`, api-key auth header); because it is synchronous, the adapter runs calls via `asyncio.to_thread`. Python 3.12 (already pinned). ## 4. Protocol annotation format (portable) Appended at the end of a template step body: ```html ``` Field rules: | field | required | notes | |-----------------|----------|----------------------------------------------------| | `resource_key` | yes | slug; stable across instances — never an item id | | `quantity` | yes | positive number consumed per execution | | `unit` | yes | must be in the unit whitelist | | `allocation` | no | `fifo` (default), `nearest_expiry`, `specific` | | `container_id` | no | only with `allocation: "specific"`; local hint, ignored if it does not match the resolved resource | | `optional` | no | `true` → skip (with a warning) if stock is missing | Parser rules (`annotate.py`): - Recognize only ``; take the **last** valid block. - Malformed JSON or schema violations ⇒ `annotation_error`, surfaced by `validate_protocol_template` and blocking step completion (never silently ignored). - `quantity` may also be `null` with `"prompt_quantity": true` for steps where the used amount varies — completion then **requires** the caller to supply quantities (voice: “how much did you use?”), otherwise it fails with a clarification request. - Visible step text is never modified by the server. Unit whitelist (extensible in config): `μL, mL, L, mg, g, kg, μg`, `ea` (each). Conversions only within the same dimension and only tested pairs (e.g. `mL↔L`, `mg↔g↔kg↔μg`); anything else ⇒ clarification, never a guess. ## 5. Resource resolution (`resolve.py`) `resource_key` → eLabFTW `items` id, resolved in this order: 1. **Local mapping store** (SQLite table): `resource_key → item_id`, written during setup. Authoritative at execution time. 2. **Resource marker** — hidden comment in the resource body: ``, auto-discovered by scanning. 3. **Configured matchers** at setup time only (CAS extra field, `custom_id`). 4. **Exact title match** — setup-time suggestion only, requires admin approval; never applied silently during execution. Ambiguous or missing mapping ⇒ execution stops with a voice-friendly clarification listing candidate resources. No fuzzy matching at runtime, ever. ### Installation on an existing instance (setup workflow) - `scan_instance` — walk `GET /items` (paginated), inventory containers, detect markers, propose matches from title/CAS/custom_id. - `map_resource` — bind `resource_key → item_id` (idempotent, upsert). - `export_mapping` / `import_mapping` — JSON mapping file for porting between instances; import proposes but still requires approval of ambiguous matches. Mapping file example: ```json { "ethanol_absolute": {"cas": "64-17-5", "title": "Absolute Ethanol", "unit": "mL"} } ``` ## 6. MCP tool surface Ten tools total. The read tools are dual-purpose: the mobile app calls them via the REST mirror (§8) to render state with zero model tokens, while the edge LLM — which receives only a minimal handoff context of ids from the app — calls the same reads over MCP when it needs detail (e.g. `get_next_protocol_step` for the full step body and stock preview). Reads (safe, any API key): | tool | purpose | |-------------------------|--------------------------------------------------------| | `find_experiments` | search by title/text/tag/custom id, compact results | | `get_experiment_context`| metadata + unfinished steps + linked resources + recent comments, all compact | | `get_next_protocol_step`| next unfinished step: text, parsed consumables, stock preview | | `list_experiment_templates` | list/search templates: id, title, short description, tags (semantic ranking, §9) | | `validate_protocol_template` | check annotations + mappings of a template/experiment | Aggregates (mutating): | tool | purpose | |-----------------------------|-------------------------------------------------------------------------| | `complete_next_protocol_step` | the headline tool — see §7 | | `complete_protocol_step` | same, with explicit step id (for “redo step 3” voice commands) | | `record_protocol_observation` | add a comment (+ optional step body edit) without completing anything | | `create_experiment_from_template` | start a new experiment from a template (interactive flow, §9) | | `adjust_inventory` | restock / correct a container, voice: “add 500 mL to …” | Setup (admin, mutating only the local mapping store): | tool | purpose | |---------------|---------------------------------------------| | `scan_instance`, `map_resource`, `export_mapping`, `import_mapping` | §5 onboarding | Result shapes are deliberately small: every tool returns compact JSON (ids, titles, quantities, statuses) — never a raw eLabFTW entity dump. Tool descriptions are one or two short sentences (small-model friendly). Mutating tools return exactly what changed, for text-to-speech readback: ```json { "ok": true, "experiment_id": 123, "step": {"id": 9, "body": "Add ethanol", "finished": true}, "consumed": [{"resource_key": "ethanol_absolute", "container_id": 12, "amount": "2.0 mL", "remaining": "48.0 mL"}], "comment_id": 77, "next_step": {"id": 10, "body": "Incubate 30 min"} } ``` ## 7. `complete_next_protocol_step` workflow (direct execution) Input: `{experiment_id, comment?, quantities?}` — nothing else. 1. `GET /experiments/{id}` + steps; select lowest-`ordering` unfinished step. None left ⇒ explicit `protocol_complete` result. 2. Parse annotation (§4). No annotation ⇒ complete step + comment, skip inventory. 3. Resolve every `resource_key` (§5). Unresolvable ⇒ clarification result, no mutation. 4. `GET /{entity_type}/{id}/containers` per resource; filter by unit compatibility; apply allocation policy (`fifo` = lowest container id with stock, `nearest_expiry` needs an expiry extra field, configured). 5. Validate stock: total available ≥ required (unit-converted). Insufficient ⇒ clarification listing what is short; `optional` consumables are skipped with a warning. 6. Execute as a journaled saga (§10), in this order: 1. decrement each container (`PATCH .../containers/{subid}` `qty_stored`), 2. finish the step (`PATCH .../steps/{subid}` `{"action":"finish"}`), 3. post the comment (`POST .../comments`). 7. Re-read the step to verify; return compact confirmation + next step. The model makes **one tool call**; all sequencing is server-side. ## 8. Read-only REST API (app fast path) The read tools are single implementations behind two transports: `rest.py` exposes them as REST endpoints so the mobile app renders state without spending LLM tokens, while `tools.py` keeps them available over MCP for the LLM and standalone clients. Mutations are **not** exposed over REST — they exist only as MCP tools, so every mutation is journaled (§10). - Auth: static bearer token (`LABVOICE_REST_TOKEN`) — the app is a trusted client; per-device tokens are an open question (§15). - Same compact result shapes and typed errors as the MCP tools (`{"error": "clarification", "message": ...}`). | endpoint | purpose | |---------------------------------------|---------------------------------------------------------------| | `GET /api/experiments?q=&limit=` | `find_experiments` | | `GET /api/experiments/{id}/state` | full app state: title, unfinished steps, linked resources, next step + parsed consumables + stock preview | | `GET /api/experiments/{id}/next-step` | just the next step (the hot path for voice) | | `GET /api/templates?q=&limit=` | `list_experiment_templates` (§9) | The `state` response is what the app renders (title, steps, stock, etc.). What the app forwards to the LLM is deliberately **minimal** — a handoff context of ids only, since LLM compute is on-device at the edge and the token budget is tight: ```json {"experiment_id": 123, "user_id": 2, "step_id": 9, "device_id": "bench-7", "location_id": "lab-2"} ``` When the model needs more than the ids — full step body, parsed consumables, stock, comments — it calls the corresponding MCP read tool (same services that back the endpoints above). REST serves the UI; MCP serves the model; one implementation of each read. ## 9. Experiment templates: semantic search & creation Voice flow ("start a new experiment for the PCR cleanup"): 1. `list_experiment_templates(query?, limit?)` — templates as `{id, title, short_description, tags}`, ranked by semantic similarity when a query is given (best match first). 2. The LLM reads the short descriptions back; the user picks one interactively. 3. `create_experiment_from_template(template_id, title?)` — `POST /experiments` with the template id; returns `{experiment_id, title, first_step}`. eLabFTW copies the template steps, so `labvoice:v1` annotations come along and the step-completion workflow (§7) applies immediately. Short description: first paragraph of the template description, truncated (`LABVOICE_TEMPLATE_DESC_LIMIT`, default 200 chars). It doubles as readback text for the LLM and as part of the search corpus. Semantic ranking (`templates.py`): - Corpus per template: title + tags + full description. - Preferred backend: local tiny embedding model (`fastembed`, ONNX, CPU, ~30 MB, `LABVOICE_EMBED_MODEL`); template vectors computed at scan time (or lazily on first use) and cached in the SQLite database as `template_id → vector`, refreshed when templates change. - Fallback (no model configured): lexical scoring — weighted token overlap, title > tags > description. Same interface, just dumber ranking. - Only templates visible to the current API key are ever returned; no fuzzy matching on ids. Creation notes: - A single `POST /experiments` — no multi-step saga; still journaled for audit (§10). - `title` optional — eLabFTW applies the template's default title format when omitted. - A created-but-unwanted experiment is archived via eLabFTW itself; this server never deletes. ## 10. Saga journal & failure handling (`journal.py`) eLabFTW has no cross-entity transactions, so every mutating workflow runs as a journaled saga in SQLite (`LABVOICE_DB_PATH`, default `~/.labvoice/journal.sqlite`): - Each execution gets an `operation_id`; journal rows record planned actions, their status, and eLabFTW responses. - Sub-actions are idempotent (re-read before write; finishing an already-finished step is a no-op; decrement uses read-modify-write with a re-read check). - On failure of step 6.2 or 6.3: **compensate** — restore decremented quantities (PATCH back), then: - success ⇒ return `reverted` result explaining what happened; - compensation fails ⇒ mark `partial_failure` in the journal **and** post an audit comment on the experiment describing exactly what is inconsistent; result tells the user which container to check. - Journal is also the audit log (who/when/what) and powers a future reconciliation tool. ## 11. Configuration & deployment Env vars (all via `config.py`): ```text ELABFTW_URL # https://eln.example.org (adapter sets host to this + /api/v2) ELABFTW_API_KEY # read-only works for read tools; writes need can_write ELABFTW_TIMEOUT=10 # per-request timeout seconds ELABFTW_RETRIES=2 LABVOICE_DB_PATH=~/.labvoice/journal.sqlite LABVOICE_REST_TOKEN=... # bearer token for the read-only REST API (§8) LABVOICE_TEMPLATE_DESC_LIMIT=200 LABVOICE_EMBED_MODEL=... # optional; unset ⇒ lexical template ranking LABVOICE_UNIT_WHITELIST=... # optional override LABVOICE_EXPIRY_FIELD=... # extra-field name for nearest_expiry allocation ``` - Transports: **stdio** (default; local/phone use) and **streamable-http** (`--transport http`, for a shared Raspberry-Pi-class deployment). In http mode a single uvicorn process serves both the MCP app and the read-only REST API (`rest.py`) — one deployment for app + LLM. `stateless_http` mode; no session affinity needed. - Health check = `GET /info` through the client (auth + reachability in one call). - Packaging: uv project, console script `labvoice`; single Dockerfile (optional). ## 12. eLabFTW endpoints used ```text GET /info GET /experiments (search: q, tags, limit/offset) GET /experiments/{id} POST /experiments {"template": , "title"?} (§9 creation) GET /{entity_type}/{id}/steps PATCH /{entity_type}/{id}/steps/{subid} {"action": "finish"} GET /{entity_type}/{id}/comments POST /{entity_type}/{id}/comments GET /items (scan/search) GET /items/{id} GET /{entity_type}/{id}/containers PATCH /{entity_type}/{id}/containers/{subid} {"qty_stored": ...} GET /storage_units?hierarchy=true (location names for readback) GET /experiments_templates (list/search templates, §9) GET /experiments_templates/{id} ``` `entity_type` ∈ {`experiments`, `items`} only — templates are read, never mutated, by this server (template editing happens in eLabFTW itself). The adapter (`elabftw.py`) normalizes all client errors into typed errors (`auth_error`, `permission_error`, `not_found`, `api_error`) with voice-friendly messages; never leaks the API key into results or logs. ## 13. Testing & verification - **Unit** (pytest, mocked `elabapi_python` API instances): annotation parser (valid/malformed/ last-block/prompt_quantity), resolver precedence + ambiguity, allocator (fifo, splitting across containers, unit conversion, insufficient stock), journal (idempotency, compensation, partial_failure), error normalization, template ranking (golden queries for semantic + lexical fallback, description truncation), creation from template (steps + annotations copied). - **REST API**: FastAPI TestClient — bearer auth enforced, `state`/`next-step` responses (snapshot tests). - **MCP-level**: tool schema lint (small input schemas), `tools/list` snapshot test, golden results for the headline workflow. - **Integration** (optional, manual): disposable eLabFTW docker instance; script creates experiment + template with annotations, maps resources, runs the saga, asserts final state via the API. - **Safety checks as tests**: read-only key must fail cleanly on every mutating tool; immutable step → explicit error; double execution of same operation_id is idempotent. ## 14. Build order (PR-sized milestones) 1. `config.py`, `errors.py`, `models.py`, elabftw adapter over `elabapi_python` + schemas (read paths only) 2. `annotate.py` parser + `validate_protocol_template` tool 3. `resolve.py` mapping store + `scan_instance` / `map_resource` tools 4. read tools: `find_experiments`, `get_experiment_context`, `get_next_protocol_step` 5. `rest.py`: read-only REST mirror over the same services (bearer auth) 6. `journal.py` saga engine 7. `workflows/protocol.py`: `complete_protocol_step` → `complete_next_protocol_step` → `record_protocol_observation` 8. `templates.py`: `list_experiment_templates` (search index) + `create_experiment_from_template` 9. `adjust_inventory` 10. MCP server assembly (stdio + http, REST mount), entry point, README, Dockerfile 11. integration script + docs (annotation authoring guide for template editors) Milestones 1–5 are read-only and independently testable; the first mutating code lands in milestone 7 on top of the journal (6). ## 15. Open questions (need your input before/at build time) 1. **Step selection** — strictly lowest `ordering` among unfinished steps, or should `complete_protocol_step` also accept a step *number* spoken by the user (“done with step three”)? (Plan assumes yes: accept id or 1-based position.) 2. **`nearest_expiry`** — is an expiry extra field available in your resources, or is `fifo` + `specific` enough for v1? (Plan: ship `fifo`/`specific` first.) 3. **Journal location on phone/mobile** — default `~/.labvoice/journal.sqlite` okay, or should the SQLite file live next to the config for easy backup? 4. **MCP SDK line** — plan targets the current stable `mcp` SDK (v2 line). Pin major version at build time. 5. **Template search backend** — ship lexical ranking first and add `fastembed` embeddings only if ranking disappoints, or embed from day one? (Plan: lexical first, same interface for both.) 6. **REST auth** — one shared bearer token per deployment enough, or per-device tokens for app installs?