Files
Labvoice/planning/mcp-server-plan.md
2026-08-30 20:44:00 +02:00

23 KiB
Raw Permalink Blame History

LabVoice MCP Server — Implementation Plan

Status: draft for review (nothing below is built yet; src/ is empty and vendor/elabapi-python holds the vendored official elabapi_python client)

1. Goal

An MCP server exposing the eLabFTW REST API v2 as a small set of aggregate tools designed for small/low-power models and voice-driven wetlab use.

Design decisions already made:

  1. Consumables live in the protocol template — machine-readable JSON at the end of the step text, hidden from the client UI in an HTML comment.
  2. Direct execution — no preview/confirmation round. The point is hands-free voice operation.
  3. Portable — templates reference stable semantic resource keys, never eLabFTW numeric IDs. The server installs on top of any existing eLabFTW instance with only template edits (plus optional resource markers).
  4. Reads are dual-purpose — the same read tools are exposed both as MCP tools and over a plain read-only REST API. REST lets the mobile app render state without spending model tokens; MCP lets the LLM fetch detail on demand. The edge LLM receives only a minimal handoff context from the app (experiment id, user id, current step id, device id, location id — on-device compute means a tight token budget), so it pulls anything beyond the ids via the read tools.
  5. Templates are first-class — new experiments are created from a template, chosen interactively via a template-listing/search tool (short descriptions, semantic ranking).

2. Architecture

      mobile app ──REST: reads for the UI (no LLM)──┐
          │                                         │
          │ voice + minimal handoff                  │
          │ context (ids only)                       │
          ▼                                         │
      small LLM ──MCP: reads on demand, aggregates──┤
          (stdio / streamable HTTP)                 │
                                                    ▼
                                   ┌──────────────────────────────┐
                                   │  mcp_server                  │
                                   │  ├── tools.py                │  small tool surface, compact results
                                   │  ├── rest.py                 │  read-only REST mirror of the reads
                                   │  ├── templates.py            │  template search index + creation
                                   │  ├── workflows/              │  aggregate operations (the point)
                                   │  │   ├── protocol.py         │  step completion saga
                                   │  │   └── inventory.py        │  allocation, unit handling
                                   │  ├── resolve.py              │  resource_key → item id mapping
                                   │  ├── annotate.py             │  annotation parse/validate
                                   │  ├── elabftw.py              │  async adapter over vendored elabapi_python
                                   │  └── journal.py              │  SQLite operation journal (sagas)
                                   └──────────────────────────────┘
                                                │  HTTPS, Authorization: <api key>
                                                ▼
                                         eLabFTW REST API v2
                                        (client: vendor/elabapi-python)

Layers, top to bottom: MCP tools / REST routes → workflow services → elabapi_python adapter → eLabFTW. The read tools are single implementations behind two transports: tools.py (MCP, for the LLM) and rest.py (REST, for the app UI). No generic PATCH/POST/DELETE passthrough tools — aggregate tools only, so a small model cannot pick an unsafe raw mutation.

3. Module layout (src/mcp_server/)

src/mcp_server/
  __init__.py
  __main__.py            # entry point: stdio (default) or streamable-http
  config.py              # env-driven settings (pydantic-settings)
  models.py              # tool input/output schemas, compact result types
  errors.py              # typed errors → voice-friendly messages
  elabftw.py             # async adapter over the vendored elabapi_python client:
                         # Configuration/ApiClient setup, error normalization,
                         # sync calls wrapped with asyncio.to_thread
  schemas.py             # pydantic models for the eLabFTW entities we use
  annotate.py            # labvoice:v1 annotation parser + validator
  resolve.py             # resource key resolution (mapping store + markers)
  allocate.py            # container selection, unit whitelist/conversion
  templates.py           # experiment templates: description extraction, semantic
                         # ranking (embedding cache in SQLite), creation flow
  journal.py             # SQLite saga journal (idempotency + compensation)
  workflows/
    __init__.py
    protocol.py          # complete_next/complete/observation workflows
    inventory.py         # adjust_inventory, stock reads
    setup.py             # scan_instance, map_resource, validate_templates
  tools.py               # MCP tool registration (one function per tool)
  rest.py                # FastAPI: read-only REST mirror of the read tools
                         # (state rendering for the app UI; bearer-token protected)
  server.py              # FastMCP assembly, lifespan; mounts rest.py in http mode
vendor/elabapi-python/   # official generated client (package elabapi_python,
                         # sync urllib3) — vendored, not edited; imported as-is

Dependencies to add: mcp (official Python SDK), pydantic-settings, aiosqlite (journal), fastapi + uvicorn (REST fast path), fastembed (optional — semantic template ranking, unset ⇒ lexical fallback). The REST client itself comes from the vendored elabapi_python package (elabapi_python.Configuration with host = ELABFTW_URL + /api/v2, api-key auth header); because it is synchronous, the adapter runs calls via asyncio.to_thread. Python 3.12 (already pinned).

4. Protocol annotation format (portable)

Appended at the end of a template step body:

<!-- labvoice:v1
{
  "consumables": [
    {"resource_key": "ethanol_absolute", "quantity": 2.0, "unit": "mL",
     "allocation": "fifo"}
  ]
}
-->

Field rules:

field required notes
resource_key yes slug; stable across instances — never an item id
quantity yes positive number consumed per execution
unit yes must be in the unit whitelist
allocation no fifo (default), nearest_expiry, specific
container_id no only with allocation: "specific"; local hint, ignored if it does not match the resolved resource
optional no true → skip (with a warning) if stock is missing

Parser rules (annotate.py):

  • Recognize only <!-- labvoice:v1 ... -->; take the last valid block.
  • Malformed JSON or schema violations ⇒ annotation_error, surfaced by validate_protocol_template and blocking step completion (never silently ignored).
  • quantity may also be null with "prompt_quantity": true for steps where the used amount varies — completion then requires the caller to supply quantities (voice: “how much did you use?”), otherwise it fails with a clarification request.
  • Visible step text is never modified by the server.

Unit whitelist (extensible in config): μL, mL, L, mg, g, kg, μg, ea (each). Conversions only within the same dimension and only tested pairs (e.g. mL↔L, mg↔g↔kg↔μg); anything else ⇒ clarification, never a guess.

5. Resource resolution (resolve.py)

resource_key → eLabFTW items id, resolved in this order:

  1. Local mapping store (SQLite table): resource_key → item_id, written during setup. Authoritative at execution time.
  2. Resource marker — hidden comment in the resource body: <!-- labvoice:resource-key=ethanol_absolute -->, auto-discovered by scanning.
  3. Configured matchers at setup time only (CAS extra field, custom_id).
  4. Exact title match — setup-time suggestion only, requires admin approval; never applied silently during execution.

Ambiguous or missing mapping ⇒ execution stops with a voice-friendly clarification listing candidate resources. No fuzzy matching at runtime, ever.

Installation on an existing instance (setup workflow)

  • scan_instance — walk GET /items (paginated), inventory containers, detect markers, propose matches from title/CAS/custom_id.
  • map_resource — bind resource_key → item_id (idempotent, upsert).
  • export_mapping / import_mapping — JSON mapping file for porting between instances; import proposes but still requires approval of ambiguous matches.

Mapping file example:

{
  "ethanol_absolute": {"cas": "64-17-5", "title": "Absolute Ethanol", "unit": "mL"}
}

6. MCP tool surface

Ten tools total. The read tools are dual-purpose: the mobile app calls them via the REST mirror (§8) to render state with zero model tokens, while the edge LLM — which receives only a minimal handoff context of ids from the app — calls the same reads over MCP when it needs detail (e.g. get_next_protocol_step for the full step body and stock preview).

Reads (safe, any API key):

tool purpose
find_experiments search by title/text/tag/custom id, compact results
get_experiment_context metadata + unfinished steps + linked resources + recent comments, all compact
get_next_protocol_step next unfinished step: text, parsed consumables, stock preview
list_experiment_templates list/search templates: id, title, short description, tags (semantic ranking, §9)
validate_protocol_template check annotations + mappings of a template/experiment

Aggregates (mutating):

tool purpose
complete_next_protocol_step the headline tool — see §7
complete_protocol_step same, with explicit step id (for “redo step 3” voice commands)
record_protocol_observation add a comment (+ optional step body edit) without completing anything
create_experiment_from_template start a new experiment from a template (interactive flow, §9)
adjust_inventory restock / correct a container, voice: “add 500 mL to …”

Setup (admin, mutating only the local mapping store):

tool purpose
scan_instance, map_resource, export_mapping, import_mapping §5 onboarding

Result shapes are deliberately small: every tool returns compact JSON (ids, titles, quantities, statuses) — never a raw eLabFTW entity dump. Tool descriptions are one or two short sentences (small-model friendly). Mutating tools return exactly what changed, for text-to-speech readback:

{
  "ok": true,
  "experiment_id": 123,
  "step": {"id": 9, "body": "Add ethanol", "finished": true},
  "consumed": [{"resource_key": "ethanol_absolute", "container_id": 12,
                "amount": "2.0 mL", "remaining": "48.0 mL"}],
  "comment_id": 77,
  "next_step": {"id": 10, "body": "Incubate 30 min"}
}

7. complete_next_protocol_step workflow (direct execution)

Input: {experiment_id, comment?, quantities?} — nothing else.

  1. GET /experiments/{id} + steps; select lowest-ordering unfinished step. None left ⇒ explicit protocol_complete result.
  2. Parse annotation (§4). No annotation ⇒ complete step + comment, skip inventory.
  3. Resolve every resource_key (§5). Unresolvable ⇒ clarification result, no mutation.
  4. GET /{entity_type}/{id}/containers per resource; filter by unit compatibility; apply allocation policy (fifo = lowest container id with stock, nearest_expiry needs an expiry extra field, configured).
  5. Validate stock: total available ≥ required (unit-converted). Insufficient ⇒ clarification listing what is short; optional consumables are skipped with a warning.
  6. Execute as a journaled saga (§10), in this order:
    1. decrement each container (PATCH .../containers/{subid} qty_stored),
    2. finish the step (PATCH .../steps/{subid} {"action":"finish"}),
    3. post the comment (POST .../comments).
  7. Re-read the step to verify; return compact confirmation + next step.

The model makes one tool call; all sequencing is server-side.

8. Read-only REST API (app fast path)

The read tools are single implementations behind two transports: rest.py exposes them as REST endpoints so the mobile app renders state without spending LLM tokens, while tools.py keeps them available over MCP for the LLM and standalone clients. Mutations are not exposed over REST — they exist only as MCP tools, so every mutation is journaled (§10).

  • Auth: static bearer token (LABVOICE_REST_TOKEN) — the app is a trusted client; per-device tokens are an open question (§15).
  • Same compact result shapes and typed errors as the MCP tools ({"error": "clarification", "message": ...}).
endpoint purpose
GET /api/experiments?q=&limit= find_experiments
GET /api/experiments/{id}/state full app state: title, unfinished steps, linked resources, next step + parsed consumables + stock preview
GET /api/experiments/{id}/next-step just the next step (the hot path for voice)
GET /api/templates?q=&limit= list_experiment_templates (§9)

The state response is what the app renders (title, steps, stock, etc.). What the app forwards to the LLM is deliberately minimal — a handoff context of ids only, since LLM compute is on-device at the edge and the token budget is tight:

{"experiment_id": 123, "user_id": 2, "step_id": 9,
 "device_id": "bench-7", "location_id": "lab-2"}

When the model needs more than the ids — full step body, parsed consumables, stock, comments — it calls the corresponding MCP read tool (same services that back the endpoints above). REST serves the UI; MCP serves the model; one implementation of each read.

9. Experiment templates: semantic search & creation

Voice flow ("start a new experiment for the PCR cleanup"):

  1. list_experiment_templates(query?, limit?) — templates as {id, title, short_description, tags}, ranked by semantic similarity when a query is given (best match first).
  2. The LLM reads the short descriptions back; the user picks one interactively.
  3. create_experiment_from_template(template_id, title?) — POST /experiments with the template id; returns {experiment_id, title, first_step}. eLabFTW copies the template steps, so labvoice:v1 annotations come along and the step-completion workflow (§7) applies immediately.

Short description: first paragraph of the template description, truncated (LABVOICE_TEMPLATE_DESC_LIMIT, default 200 chars). It doubles as readback text for the LLM and as part of the search corpus.

Semantic ranking (templates.py):

  • Corpus per template: title + tags + full description.
  • Preferred backend: local tiny embedding model (fastembed, ONNX, CPU, ~30 MB, LABVOICE_EMBED_MODEL); template vectors computed at scan time (or lazily on first use) and cached in the SQLite database as template_id → vector, refreshed when templates change.
  • Fallback (no model configured): lexical scoring — weighted token overlap, title > tags > description. Same interface, just dumber ranking.
  • Only templates visible to the current API key are ever returned; no fuzzy matching on ids.

Creation notes:

  • A single POST /experiments — no multi-step saga; still journaled for audit (§10).
  • title optional — eLabFTW applies the template's default title format when omitted.
  • A created-but-unwanted experiment is archived via eLabFTW itself; this server never deletes.

10. Saga journal & failure handling (journal.py)

eLabFTW has no cross-entity transactions, so every mutating workflow runs as a journaled saga in SQLite (LABVOICE_DB_PATH, default ~/.labvoice/journal.sqlite):

  • Each execution gets an operation_id; journal rows record planned actions, their status, and eLabFTW responses.
  • Sub-actions are idempotent (re-read before write; finishing an already-finished step is a no-op; decrement uses read-modify-write with a re-read check).
  • On failure of step 6.2 or 6.3: compensate — restore decremented quantities (PATCH back), then:
    • success ⇒ return reverted result explaining what happened;
    • compensation fails ⇒ mark partial_failure in the journal and post an audit comment on the experiment describing exactly what is inconsistent; result tells the user which container to check.
  • Journal is also the audit log (who/when/what) and powers a future reconciliation tool.

11. Configuration & deployment

Env vars (all via config.py):

ELABFTW_URL                    # https://eln.example.org  (adapter sets host to this + /api/v2)
ELABFTW_API_KEY                # read-only works for read tools; writes need can_write
ELABFTW_TIMEOUT=10             # per-request timeout seconds
ELABFTW_RETRIES=2
LABVOICE_DB_PATH=~/.labvoice/journal.sqlite
LABVOICE_REST_TOKEN=...         # bearer token for the read-only REST API (§8)
LABVOICE_TEMPLATE_DESC_LIMIT=200
LABVOICE_EMBED_MODEL=...        # optional; unset ⇒ lexical template ranking
LABVOICE_UNIT_WHITELIST=...    # optional override
LABVOICE_EXPIRY_FIELD=...      # extra-field name for nearest_expiry allocation
  • Transports: stdio (default; local/phone use) and streamable-http (--transport http, for a shared Raspberry-Pi-class deployment). In http mode a single uvicorn process serves both the MCP app and the read-only REST API (rest.py) — one deployment for app + LLM. stateless_http mode; no session affinity needed.
  • Health check = GET /info through the client (auth + reachability in one call).
  • Packaging: uv project, console script labvoice; single Dockerfile (optional).

12. eLabFTW endpoints used

GET    /info
GET    /experiments                      (search: q, tags, limit/offset)
GET    /experiments/{id}
POST   /experiments                      {"template": <id>, "title"?}   (§9 creation)
GET    /{entity_type}/{id}/steps
PATCH  /{entity_type}/{id}/steps/{subid}      {"action": "finish"}
GET    /{entity_type}/{id}/comments
POST   /{entity_type}/{id}/comments
GET    /items                            (scan/search)
GET    /items/{id}
GET    /{entity_type}/{id}/containers
PATCH  /{entity_type}/{id}/containers/{subid} {"qty_stored": ...}
GET    /storage_units?hierarchy=true     (location names for readback)
GET    /experiments_templates            (list/search templates, §9)
GET    /experiments_templates/{id}

entity_type ∈ {experiments, items} only — templates are read, never mutated, by this server (template editing happens in eLabFTW itself). The adapter (elabftw.py) normalizes all client errors into typed errors (auth_error, permission_error, not_found, api_error) with voice-friendly messages; never leaks the API key into results or logs.

13. Testing & verification

  • Unit (pytest, mocked elabapi_python API instances): annotation parser (valid/malformed/ last-block/prompt_quantity), resolver precedence + ambiguity, allocator (fifo, splitting across containers, unit conversion, insufficient stock), journal (idempotency, compensation, partial_failure), error normalization, template ranking (golden queries for semantic + lexical fallback, description truncation), creation from template (steps + annotations copied).
  • REST API: FastAPI TestClient — bearer auth enforced, state/next-step responses (snapshot tests).
  • MCP-level: tool schema lint (small input schemas), tools/list snapshot test, golden results for the headline workflow.
  • Integration (optional, manual): disposable eLabFTW docker instance; script creates experiment + template with annotations, maps resources, runs the saga, asserts final state via the API.
  • Safety checks as tests: read-only key must fail cleanly on every mutating tool; immutable step → explicit error; double execution of same operation_id is idempotent.

14. Build order (PR-sized milestones)

  1. config.py, errors.py, models.py, elabftw adapter over elabapi_python + schemas (read paths only)
  2. annotate.py parser + validate_protocol_template tool
  3. resolve.py mapping store + scan_instance / map_resource tools
  4. read tools: find_experiments, get_experiment_context, get_next_protocol_step
  5. rest.py: read-only REST mirror over the same services (bearer auth)
  6. journal.py saga engine
  7. workflows/protocol.py: complete_protocol_step → complete_next_protocol_step → record_protocol_observation
  8. templates.py: list_experiment_templates (search index) + create_experiment_from_template
  9. adjust_inventory
  10. MCP server assembly (stdio + http, REST mount), entry point, README, Dockerfile
  11. integration script + docs (annotation authoring guide for template editors)

Milestones 1–5 are read-only and independently testable; the first mutating code lands in milestone 7 on top of the journal (6).

15. Open questions (need your input before/at build time)

  1. Step selection — strictly lowest ordering among unfinished steps, or should complete_protocol_step also accept a step number spoken by the user (“done with step three”)? (Plan assumes yes: accept id or 1-based position.)
  2. nearest_expiry — is an expiry extra field available in your resources, or is fifo + specific enough for v1? (Plan: ship fifo/specific first.)
  3. Journal location on phone/mobile — default ~/.labvoice/journal.sqlite okay, or should the SQLite file live next to the config for easy backup?
  4. MCP SDK line — plan targets the current stable mcp SDK (v2 line). Pin major version at build time.
  5. Template search backend — ship lexical ranking first and add fastembed embeddings only if ranking disappoints, or embed from day one? (Plan: lexical first, same interface for both.)
  6. REST auth — one shared bearer token per deployment enough, or per-device tokens for app installs?