Docs
Ineffable serves Laya decision models: you send a text and a set of typed questions, and get back an answer with a probability for every option. There is one inference endpoint and a management API; the dashboard, the command line tool and the MCP server all use the same API.
Quickstart
Create an API key in the dashboard under API keys, then:
curl https://api.ineffable.nl/v1/systemone \
-H "Authorization: Bearer $INEFFABLE_API_KEY" \
-d '{
"model": "laya",
"state": {"body": "billed twice, refund please or we cancel"},
"questions": {
"dept": {"type": "choice", "instructions": "Which team?",
"criteria": {"billing": "refunds and invoices", "tech": "bugs and outages"}},
"urgency": {"type": "score", "instructions": "How urgent?",
"criteria": ["not urgent", "normal", "urgent", "critical"]},
"churn": {"type": "noul", "instructions": "The customer threatens to leave."}
}
}'
The endpoint speaks the TypeSafe Jev and laya-serve wire format, so existing clients work by changing their base URL.
Questions
state is a string, a JSON object or a list of conversation turns. questions maps your own question ids to questions of three types.
| Type | criteria | Answer |
|---|---|---|
choice | An object of label to description, or a list of labels | the chosen label |
score | A list of level descriptions, lowest first; describe every level | the expected level, a number |
noul | Optional {"false": "…", "true": "…"} | the probability that the statement holds |
Describe options the way you would explain them to a new colleague. Descriptions matter more than labels.
Answers
{"model": "nhsd/tickets@v3",
"answers": {
"dept": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.9789, "tech": 0.0211},
"answer_confidence": 0.9789, "confidence": 0.8524,
"action": {"act_probability": 1.0}}},
"usage": {"input_tokens": 153, "output_tokens": 0}}
answer_confidence is the calibrated probability of the answer given, on every question type. Compare it with a threshold to decide whether to act or escalate; each fine-tune report has a table that shows, per threshold, how many answers the model takes on and how often it is right on those. Response headers carry X-Decision-Id, X-Model-Checkpoint and X-Inference-Time-Ms.
Model names
| Name | Means |
|---|---|
laya | The stock English checkpoint |
laya-multilingual | The stock multilingual checkpoint |
laya-typed-decisions | Stock, tuned on four operational workflows |
omitted, or auto | Picks English or multilingual from the text's script and language |
acme/router | Your model, its live version |
acme/router@v3 | A specific version |
Training data
JSONL, one record per line. gold holds the right answers: a choice label, a score level index, or true/false. Soft labels via probabilities are optional and help when people disagree.
{"state": {"ticket": "I was charged twice this month"},
"questions": {"dept": {"type": "choice", "instructions": "Which team?",
"criteria": {"billing": "refunds", "tech": "bugs", "sales": "plans"}}},
"gold": {"dept": {"label": "billing"}}}
Validation runs before anything is stored and reports line numbers. It rejects datasets with fewer than 50 labelled sequences, choice questions with more than 20 options, or more than 5% of sequences over 1,024 tokens; it warns about rare or dominant labels and conflicting duplicates.
Fine-tuning
ineffable dataset upload tickets.jsonl --name "support tickets"
ineffable finetune --dataset <id> --base laya --model acme/router --preset fast --deploy
ineffable job wait <job id>
A job holds back 10% of the data, trains, fits calibration on the held-out part and reports accuracy, calibration error and the thresholds table, compared with the base model and with the version that is live. Presets: official is the published Laya recipe; fast batches sequences of similar length and trains about 1.4 times faster at the same accuracy. Every setting can be overridden with --set key=value; ineffable finetune recipe lists them with their ranges.
A checkpoint trained elsewhere can be uploaded as a .tar.gz with ineffable checkpoint upload, and any checkpoint can be downloaded again.
Versions and deploys
Each fine-tune or upload registered to a model becomes its next version. Deploying a version switches traffic on the next request; requests already running finish on the version they started with. Roll back with one call:
ineffable models deploy acme/router --version 3
ineffable models rollback acme/router
Decisions and labels
Every answer is logged with its request, unless your organisation turns request storage off. Filter by model, outcome or label; label a decision with the right answer; export labelled decisions as JSONL and they are a training set as they stand:
ineffable decisions export next-round.jsonl --model acme/router --labelled
MCP and CLI
The MCP server gives an agent such as Claude every step above: validating data, starting and waiting on fine-tunes, reading reports, deploying and asking questions. It is hosted at https://api.ineffable.nl/mcp. Add it and sign in when Claude asks; you approve the connection in the dashboard, and it appears on the Keys page, where you can revoke it:
claude mcp add --transport http ineffable https://api.ineffable.nl/mcp
On claude.ai, add it as a custom connector with the same URL. For scripts, skip the sign-in and pass an API key with the infer and manage scopes: add --header "Authorization: Bearer inf_...".
The hosted server takes datasets as text, up to 16 MB. For larger files and checkpoint uploads, the ineffable package has the command line tool and a local MCP server that reads files from your machine. You get the package with your access.
export INEFFABLE_URL=https://api.ineffable.nl INEFFABLE_API_KEY=inf_...
claude mcp add ineffable -e INEFFABLE_URL -e INEFFABLE_API_KEY -- ineffable-mcp
Limits and errors
| Limit | Value |
|---|---|
| Request body | 1 MB (16 MB for /v1/systemone/batch, up to 256 states) |
| Questions per request | 64; 255 options per question, 1,024 in total |
| State | 200,000 characters; the model reads the first 1,024 tokens |
| Rate | 600 requests a minute per key, 1,800 per organisation |
Errors: 400 malformed body · 401 missing or unknown key · 403 key without the scope · 404 unknown model · 413 over a limit · 422 a question the model cannot read (the message says which and why) · 429 over the rate, retry after the Retry-After seconds · 503 the model is loading or the GPU is training, retry after Retry-After. Management endpoints under /api/v1 answer errors as {"error": {"code", "message", "details"}}; the OpenAPI description is at api.ineffable.nl/docs.