Docs

Ineffable serves Laya decision models: you send a text and a set of typed questions, and get back an answer with a probability for every option. There is one inference endpoint and a management API; the dashboard, the command line tool and the MCP server all use the same API.

Quickstart

Create an API key in the dashboard under API keys, then:

curl https://api.ineffable.nl/v1/systemone \
  -H "Authorization: Bearer $INEFFABLE_API_KEY" \
  -d '{
    "model": "laya",
    "state": {"body": "billed twice, refund please or we cancel"},
    "questions": {
      "dept":    {"type": "choice", "instructions": "Which team?",
                  "criteria": {"billing": "refunds and invoices", "tech": "bugs and outages"}},
      "urgency": {"type": "score", "instructions": "How urgent?",
                  "criteria": ["not urgent", "normal", "urgent", "critical"]},
      "churn":   {"type": "noul", "instructions": "The customer threatens to leave."}
    }
  }'

The endpoint speaks the TypeSafe Jev and laya-serve wire format, so existing clients work by changing their base URL.

Questions

state is a string, a JSON object or a list of conversation turns. questions maps your own question ids to questions of three types.

TypecriteriaAnswer
choiceAn object of label to description, or a list of labelsthe chosen label
scoreA list of level descriptions, lowest first; describe every levelthe expected level, a number
noulOptional {"false": "…", "true": "…"}the probability that the statement holds

Describe options the way you would explain them to a new colleague. Descriptions matter more than labels.

Answers

{"model": "nhsd/tickets@v3",
 "answers": {
   "dept": {"type": "choice", "choice": "billing",
            "probabilities": {"billing": 0.9789, "tech": 0.0211},
            "answer_confidence": 0.9789, "confidence": 0.8524,
            "action": {"act_probability": 1.0}}},
 "usage": {"input_tokens": 153, "output_tokens": 0}}

answer_confidence is the calibrated probability of the answer given, on every question type. Compare it with a threshold to decide whether to act or escalate; each fine-tune report has a table that shows, per threshold, how many answers the model takes on and how often it is right on those. Response headers carry X-Decision-Id, X-Model-Checkpoint and X-Inference-Time-Ms.

Model names

NameMeans
layaThe stock English checkpoint
laya-multilingualThe stock multilingual checkpoint
laya-typed-decisionsStock, tuned on four operational workflows
omitted, or autoPicks English or multilingual from the text's script and language
acme/routerYour model, its live version
acme/router@v3A specific version

Training data

JSONL, one record per line. gold holds the right answers: a choice label, a score level index, or true/false. Soft labels via probabilities are optional and help when people disagree.

{"state": {"ticket": "I was charged twice this month"},
 "questions": {"dept": {"type": "choice", "instructions": "Which team?",
                        "criteria": {"billing": "refunds", "tech": "bugs", "sales": "plans"}}},
 "gold": {"dept": {"label": "billing"}}}

Validation runs before anything is stored and reports line numbers. It rejects datasets with fewer than 50 labelled sequences, choice questions with more than 20 options, or more than 5% of sequences over 1,024 tokens; it warns about rare or dominant labels and conflicting duplicates.

Fine-tuning

ineffable dataset upload tickets.jsonl --name "support tickets"
ineffable finetune --dataset <id> --base laya --model acme/router --preset fast --deploy
ineffable job wait <job id>

A job holds back 10% of the data, trains, fits calibration on the held-out part and reports accuracy, calibration error and the thresholds table, compared with the base model and with the version that is live. Presets: official is the published Laya recipe; fast batches sequences of similar length and trains about 1.4 times faster at the same accuracy. Every setting can be overridden with --set key=value; ineffable finetune recipe lists them with their ranges.

A checkpoint trained elsewhere can be uploaded as a .tar.gz with ineffable checkpoint upload, and any checkpoint can be downloaded again.

Versions and deploys

Each fine-tune or upload registered to a model becomes its next version. Deploying a version switches traffic on the next request; requests already running finish on the version they started with. Roll back with one call:

ineffable models deploy acme/router --version 3
ineffable models rollback acme/router

Decisions and labels

Every answer is logged with its request, unless your organisation turns request storage off. Filter by model, outcome or label; label a decision with the right answer; export labelled decisions as JSONL and they are a training set as they stand:

ineffable decisions export next-round.jsonl --model acme/router --labelled

MCP and CLI

The MCP server gives an agent such as Claude every step above: validating data, starting and waiting on fine-tunes, reading reports, deploying and asking questions. It is hosted at https://api.ineffable.nl/mcp. Add it and sign in when Claude asks; you approve the connection in the dashboard, and it appears on the Keys page, where you can revoke it:

claude mcp add --transport http ineffable https://api.ineffable.nl/mcp

On claude.ai, add it as a custom connector with the same URL. For scripts, skip the sign-in and pass an API key with the infer and manage scopes: add --header "Authorization: Bearer inf_...".

The hosted server takes datasets as text, up to 16 MB. For larger files and checkpoint uploads, the ineffable package has the command line tool and a local MCP server that reads files from your machine. You get the package with your access.

export INEFFABLE_URL=https://api.ineffable.nl INEFFABLE_API_KEY=inf_...
claude mcp add ineffable -e INEFFABLE_URL -e INEFFABLE_API_KEY -- ineffable-mcp

Limits and errors

LimitValue
Request body1 MB (16 MB for /v1/systemone/batch, up to 256 states)
Questions per request64; 255 options per question, 1,024 in total
State200,000 characters; the model reads the first 1,024 tokens
Rate600 requests a minute per key, 1,800 per organisation

Errors: 400 malformed body · 401 missing or unknown key · 403 key without the scope · 404 unknown model · 413 over a limit · 422 a question the model cannot read (the message says which and why) · 429 over the rate, retry after the Retry-After seconds · 503 the model is loading or the GPU is training, retry after Retry-After. Management endpoints under /api/v1 answer errors as {"error": {"code", "message", "details"}}; the OpenAPI description is at api.ineffable.nl/docs.