Conscious Loop API

Conscious Loop API

Pipelines, samples, feedback, versions and benchmarks. What each thing is and how it behaves is in the Conscious Loop guide; this page lists every route with a request and a response.

Authentication & Errors

The base URL is https://api.runbios.ai. Send your API key as X-API-Key: reads need the loop:read scope and everything else loop:write. A key for one workspace needs nothing more; a key for several also sends X-Workspace-ID. Bodies and answers are JSON. An unknown field in a pipeline body is refused with 400 naming it, and so are judge_id, judge_model, max_tokens and per_run_ceiling_cents in a benchmark body.

Errors use the platform’s envelope, {"error": {"code": "…", "message": "…"}}. Branch on code; the message is a sentence for a person. Every code the loop answers:

403 LOOP_COMING_SOON
When
Every route where the platform has the Conscious Loop switched off, except the five that let you stop paying: GET /api/loop/pipelines, GET /api/loop/pipelines/{id}, PUT /api/loop/pipelines/{id} whose body is only {"enabled": false} (with at most expected_revision), DELETE /api/loop/pipelines/{id} and POST /api/loop/training-runs/{id}/cancel, which answer as usual (and whose five tools npx runbios-mcp still lists there).
422 MODEL_NOT_TRAINABLE
When
The model is not in GET /api/loop/models or is larger than the platform trains, the challenger is the same as the model, or the challenger cannot be pinned to an exact revision.
422 NO_GPU_FITS
When
No machines can train or run the model right now. Nothing was saved.
422 NO_JUDGE_MODEL
When
Creating a pipeline when no model your workspace can call is able to score its attempts (see Automatic test). Nothing was saved.
422 MODEL_REVISION_UNAVAILABLE
When
The model could not be pinned to an exact revision because a service did not answer. Nothing was saved; try again.
422 PROMOTION_POLICY_UNKNOWN
When
promotion.policy is not one of holdout, all, primary, k_of_n, weighted, manual.
422 PROMOTION_K_REQUIRED
When
k_of_n without a k of 1 or more.
422 PROMOTION_K_UNREACHABLE
When
k is more than the measurements the pipeline takes.
422 PROMOTION_POLICY_NEEDS_BENCHMARK
When
primary with no benchmark attached.
422 METHOD_NOT_ENABLED
When
A pipeline save whose body asks for preference or reward training: training_method, rlhf_type, method or algorithm naming rlhf, dpo, kto, grpo or the like, at the top or under training. Training that learns from rankings or rewards (like DPO or GRPO) is not available yet; every version learns from your good and corrected answers. Nothing was saved.
503 TRAINING_CURVE_UNAVAILABLE
When
How a training went while it trained could not be read just then. Nothing is wrong with the training; try again in a minute.
503 GPU_OPTIONS_UNAVAILABLE
When
The machines could not be chosen because a service did not answer. Nothing was saved; try again.
503 MODELS_UNAVAILABLE
When
The models a pipeline can train could not be read just then, so the model could not be checked. Nothing was saved; try again in a minute. GET /api/loop/models answers the same while it lasts.
503 AGENT_COULD_NOT_START
When
Creating a pipeline, or Train now, when the key its attempts are scored with could not be made just then. Nothing was saved or started; try again in a minute.
409 PIPELINE_NAME_TAKEN
When
The name is used by another pipeline, including a deleted one.
409 OWN_TAG_TAKEN
When
Another pipeline already owns the tag this name makes.
409 DATASET_NAME_TAKEN
When
Saving or renaming a dataset with a name another dataset of the workspace already has, whatever its case. Choose another name.
409 REVISION_MISMATCH
When
The pipeline or dataset changed since you read it, or while your edit was being saved (for a pipeline, a new version went live). Nothing was saved. Read it again and resend with the new expected_revision.
409 NOT_A_SIMPLE_PIPELINE
When
Changing tags, conditions, datasets, use_case, challenger_model or training of a pipeline made before pipelines had their own tag, or importing into one.
409 TRAINING_NOT_ENABLED
When
Creating a pipeline, or changing its model or switching it on, while the platform's services are being updated and its training service would not take it. Nothing was saved and nothing in the request is wrong: try again in a few minutes.
422 NOT_ENOUGH_SAMPLES
When
Train now on a pipeline that is short. The body also has have and need.
422 TRAINING_SETTING_OUT_OF_RANGE
When
A value of training outside its range: epochs 1 to 15, learning_rate 0.000001 to 0.01, lora_rank 4, 8, 16, 32 or 64, lora_alpha 1 to 512, replay_percent 0 to 50. The message names the field and its range. Nothing was saved.
409 RUN_ACTIVE
When
An attempt is already running (train now), the model cannot change while one runs, or a version cannot be rolled back while an attempt is being compared with it or built on top of it (cancel that attempt first).
409 RULE_PAUSED
When
The pipeline is switched off or paused. Save it with "enabled": true to resume.
409 PIPELINE_NEEDS_SAVE
When
Train now on a pipeline whose settings changed without being saved as a pipeline since (one made before pipelines existed, or moved by a platform upgrade). Save its settings, then train again.
409 VERSION_LIMIT_REACHED
When
The pipeline has made max_versions versions. A rebuild that takes out a conversation you deleted is not held by the limit.
409 MONTHLY_LIMIT_REACHED
When
Train now when one more attempt would not fit in the pipeline’s monthly safety limit. The body also has resumes_at, the 1st of next month (UTC). The limit is set by the platform from the model’s prices, and its figures are not fields of any answer.
409 SCORING_CAPPED
When
Train now while your workspace’s Conscious Loop key is at its monthly spending cap. The body also has resumes_at.
409 LIVE_VERSION_STOPPED
When
Train now while the deployment your live version answers from is stopped, paused or gone, so an attempt could not be compared with it. Resume it in Deployments or roll the version back on the Versions tab, then train again.
409 AGENT_OFF
When
Train now when the key attempts are scored with is still unusable after it was renewed. Nothing was started; try again in a minute.
409 NOT_REVIEWABLE
When
Approving or rejecting an attempt that is not waiting for a decision.
409 INCONCLUSIVE_REQUIRES_FORCE
When
Approving an inconclusive attempt without "force": true.
409 CONSENT_CHANGED
When
Approving an attempt while the pipeline’s latest settings are waiting to be accepted: save the pipeline (its Settings, or any PUT /api/loop/pipelines/{id}) to accept them, then approve. Also when expected_revision is not the pipeline’s revision: read the pipeline again. An attempt that ran under earlier settings that have since been saved is approved under the settings as they stand, not refused.
409 CANDIDATE_MISSING
When
Approving an attempt whose trained model is no longer there.
409 NOT_PROMOTED
When
Rolling back an attempt that never went live.
409 ROLLBACK_WINDOW_CLOSED
When
More than 30 days since the version went live.
409 ROLLBACK_SUPERSEDED
When
A later version has gone live since. Roll back the most recent one.
409 NOTHING_TO_RESTORE
When
Rolling back a version that does not record what was serving before it. Nothing is changed rather than guessed.
409 ROLLBACK_DELETED_DATA
When
Rolling back when the version it would put back learned from a conversation that was deleted since, or was built on a version that did. It never answers again; a rebuild without that conversation replaces it (see Deleting a sample a version learned).
CHAIN_MOVED
When
Not a refusal: an attempt’s error_code when the version it was built on was rolled back while it trained. It is set aside with no comparison paid for, and its samples go to the next attempt.
NO_MERGED_EXPORT
When
Not a refusal: an attempt’s error_code when its training finished without the merged model it would be served as. Its samples go to the next attempt.
TRAINING_SERVICE_BEHIND
When
Not a refusal: an attempt’s error_code when it was built on the live version and the platform’s training service was being updated when it asked for its training. It stopped before any training started, nothing was charged, and the next attempt asks again. STEP_STALLED:TRAINING_SERVICE_BEHIND is the same after six hours of waiting for that update; the pipeline tries again on its own.
STEP_STALLED:SERVING_NOT_CERTIFIED
When
Not a refusal: a training’s error_code when the platform had not yet checked that it can test new versions of its model on the current serving engine. It waited six hours before its training and stood down: nothing was charged, the pipeline is not paused, and its next training waits the same way until that check is done. CHECKPOINT_NOT_DEPLOYABLE:SERVING_NOT_CERTIFIED is the same found at its test, after it trained; the next training waits before it starts.
ENV_GATED:SWITCHED_OFF
When
Not a refusal: an attempt’s error_code when the platform switched the Conscious Loop off before its training started. It stopped there, no GPU time was charged, and the pipeline trains again on its own once the loop is back.
MERGED_EXPORT_FAILED
When
Not a refusal: an attempt’s error_code is TRAINING_FAILED:MERGED_EXPORT_FAILED when it was built on the live version and its training finished but its merged model could not be saved. Its GPU time is billed as for any failed training, and its samples go to the next attempt.
400 CHAINED_ADAPTER_NOT_DEPLOYABLE
When
From the Deployments API, not this one: deploying yourself the adapter checkpoint of a version built on earlier ones. It means nothing without them; deploy its merged model (the checkpoint whose format is merged) instead.
CHECKPOINT_NOT_DEPLOYABLE:CHAIN_HEAD
When
Not a refusal: an attempt’s error_code starts so when it was built on the live version, the live deployment was not the model it was built on (stopped, or changed to another precision or context), and the platform would not start that version on a machine of its own to compare with. One whose machine could not start ends CANDIDATE_BOOT_FAILED:CHAIN_HEAD_ and the machine’s status. Nothing was compared, and its samples go to the next attempt.
409 RUN_FINISHED
When
Stopping an attempt that has already ended.
409 NOT_TESTED
When
Approving a training that was never tested, with or without force: nothing untested goes live. Test it again first.
409 CANNOT_TEST_AGAIN
When
Testing a training again when it cannot be: the message says why, for example that it was already tested, or that the amount set aside for testing it is spent.
404 VERSION_NOT_FOUND
When
Comparing versions when a or b is not one of the pipeline’s versions.
409 SAMPLE_NOT_YET_RECORDED
When
Feedback for a sample id the gateway issued in the last few minutes (X-Loop-Sample-Id) that has not been recorded yet. Nothing was stored: wait Retry-After (2 s) and send it again, so retry after a couple of seconds. After five minutes an id that never arrived is 404 NOT_FOUND.
422 BENCHMARK_SET_TOO_SMALL
When
A benchmark of fewer than 20 conversations.
422 BENCHMARK_SET_TOO_LARGE
When
A benchmark of more than 200 conversations.
422 BENCHMARK_ROW_UNUSABLE
When
A conversation named for a benchmark has nothing to ask a model, or is too large to keep in storage. Leave it out and try again.
422 BENCHMARK_ROW_HAS_NO_ANSWER
When
An exact-check benchmark names conversations with no right answer, and the message names them. Mark each answer good or write a correction, leave them out, or use the judge instead.
422 BENCHMARK_JUDGE_HAS_NO_MODEL
When
Creating a benchmark when no model your workspace can call is able to score it. Nothing was saved.
409 BENCHMARK_RETIRED
When
Adding a retired benchmark to a pipeline.
409 BENCHMARK_ALREADY_ATTACHED
When
Adding a benchmark the pipeline already has. Move it with PUT instead.
404 BENCHMARK_NOT_ATTACHED
When
Moving or removing a benchmark that is not on the pipeline’s list.
429 RATE_LIMITED
When
A search (q) of a report’s conversations while this workspace already has one running, or while the service runs as many as it can at once. Try again after Retry-After seconds.
500 IMPORT_INCOMPLETE
When
Some rows or their feedback were not stored. The body carries every count and the import_id to retry with.

Besides these, any route can answer the platform’s own codes: 400 BAD_REQUEST (the message names what to change; a pipeline body carrying train_when.min_samples is refused this way, because the minimums are fixed), 401 UNAUTHORIZED, 403 FORBIDDEN, 404 NOT_FOUND (also for an id in another workspace, and for a deleted pipeline and its attempts, reports and benchmark results), 409 CONFLICT (a benchmark name already in use, or a conversation kept in your own storage, or a report’s conversations kept in stored objects, that could not be read just then, or a new benchmark’s conversations that could not be stored just then: nothing was saved, so try again), 413 PAYLOAD_TOO_LARGE and 500 INTERNAL_ERROR.

Recording With One Header

Your app’s own model calls record samples when they carry one header: on /v1/chat/completions for serverless models and deployments alike, and on /v1/messages for serverless models (a deployment, a pipeline’s loop-<tag> model included, serves the OpenAI shape and answers /v1/messages with 400), streaming or not. The conversation is stored with the tags of the pipelines that exist, its tool calls included (OpenAI tool_calls and Anthropic tool_use blocks), and the response says which sample it became. See Recording a call, and the guide’s Connect your app.

X-Loop-Pipeline
Where
Request
Meaning
One pipeline tag, or several separated by commas. The call is recorded with those tags even when recording is off for its model: the header is the caller’s consent for that one call. A tag no pipeline has (misspelt, or a deleted pipeline’s) is not recorded under, and is counted in GET /api/loop/stats capture.
X-Loop-Capture: off
Where
Request
Meaning
Never record this call, whatever else it says.
X-Loop-Sample-Id
Where
Response
Meaning
The id of the sample the call became (a UUID), sent before the body when streaming. Give feedback on it with POST /api/loop/traces/{id}/signals: if the sample is not recorded yet (409 SAMPLE_NOT_YET_RECORDED), retry after a couple of seconds. A stream cut short, or a repeated X-Request-ID, records nothing new under it; an answer cut short or too large to keep is counted in GET /api/loop/stats capture.
bash
curl -i https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $RUNBIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Loop-Pipeline: support-replies" \
  -d '{"model": "loop-support-replies", "messages": [{"role": "user", "content": "How long do refunds take?"}]}'

# HTTP/2 200
# x-loop-sample-id: 7c2e0b84-15da-4f37-9a6e-8d13c5f0b2a7

Samples & Feedback

A sample is stored as a trace, feedback as a signal, and tags as labels.

Feedback in parallel. Each signal is its own call, so send them one by one or several at once. The SDKs’ signalMany (TypeScript) and signal_many (Python) keep up to 8 in flight (at most 32), retry 409 SAMPLE_NOT_YET_RECORDED as a single call does, and return every item in order with its signal or its error, so one refused item costs no other. The words are the ones a sample’s own feedback takes: good, bad and corrected mean accepted, rejected and edited on every door, and the answer names the stored word.

results = client.loop.signal_many([
    {"trace_id": id_a, "verdict": "good"},
    {"trace_id": id_b, "verdict": "corrected", "correction": "Refunds take five working days."},
], concurrency=8)
failed = [r for r in results if r.get("error")]

Import

Import JSONL (one JSON object per line) or a JSON array of objects, at most 5,000 rows per call: POST /api/loop/pipelines/{id}/import adds the pipeline’s own tag to every row, and POST /api/loop/import takes the same body without a pipeline (then source is required and labels are the tags). The console sends a bigger file in pages of 5,000 under one import_id, so importing it again after a failure adds nothing twice.

Send each row in the shape it already has. The shape decides what it becomes:

{"messages": [{"role": "user", …}, {"role": "assistant", …}]}
Becomes
An answer with no feedback. It waits for review, or is good with default_feedback. Every turn is kept, system and developer messages and tool calls with their results included; the last assistant turn is the answer.
{"prompt": "…", "completion": "…"}
Becomes
The same, in the flat form.
{"messages": […], "label": true}
Becomes
Good (true) or bad (false), beside messages or beside prompt and completion.
{"prompt": "…", "completion": "…", "correction": "…"}
Becomes
A correction: the answer was wrong, and correction (or ground_truth beside a completion) is what it should have been, and what is learned from.
{"prompt": "…", "chosen": "…", "rejected": "…"}
Becomes
A correction: chosen is learned from.
{"prompt": "…", "ground_truth": "…"}
Becomes
A question with the answer it should get: ground_truth is recorded as a correction, learned from and used to judge attempts.
{"prompt": "…"}
Becomes
A question with no answer. It is kept, but there is nothing to learn from until it has one.
{"text": "…"}
Becomes
Refused: it is not a conversation.
{"…", "feedback": "bad"}
Becomes
Refused: feedback and verdict are not read, so the row is skipped rather than stored with no feedback. Use label or correction.
json
{"messages": [{"role": "user", "content": "Where is my order?"}, {"role": "assistant", "content": "It left the warehouse this morning."}]}
{"prompt": "Can I change the delivery address?", "completion": "Yes, until the order ships."}
{"messages": [{"role": "user", "content": "Do you ship abroad?"}, {"role": "assistant", "content": "No."}], "label": false}
{"prompt": "Why was I charged twice?", "chosen": "I have refunded the duplicate today.", "rejected": "Please see our billing policy."}

default_feedback (pipeline import only): "good" records every answered row that has no feedback of its own as a good example to learn from. Leave it out and those rows wait for review. Rows that bring their own label, correction or chosen/rejected keep it either way, and sending a file again never overrules feedback somebody has given a sample since. Feedback that arrives with a file is recorded as coming from the import, and every sample records whether it was captured or imported. labels are checked before anything is stored: one that cannot be a tag is refused by name, and a pipeline import counts its own tag among the 32 a sample can carry.

The answer describes what this call did, not what the pipeline holds:

imported
Meaning
Samples this call created. trace_ids names every sample from the file that is now stored.
already_present
Meaning
Rows an earlier call you named as the same import already stored (by import_id or row_ids), so they were not stored twice.
reviewed
Meaning
The verdicts this call wrote. It can be non-zero on a retry, when a verdict that failed the first time is written now.
needs_review
Meaning
Samples this call created that have no feedback.
refused
Meaning
Rows that could not be read as a conversation, with reasons in refused_why. Fix the file.
not_saved
Meaning
Rows understood and then not stored (not_saved_rows gives their positions, from 1). The file is fine: retry.
verdicts_not_saved
Meaning
Rows stored without the feedback they came with. Retry to write it.
by_shape
Meaning
What this call created, per row shape.
import_id
Meaning
The token this call was filed under. Send it again to retry the call.

A call with any not_saved or verdicts_not_saved answers with the code IMPORT_INCOMPLETE and all of the counts above. To retry, send the same rows with the same import_id. Rows already stored come back under already_present, and reviewed counts the verdicts this call wrote, so it confirms that a repair landed. A call with no import_id is a new import and stores everything again.

For files over 5,000 rows, give every row its own id in row_ids (a ticket number, say): then pages, order and subsets stop mattering. With row_ids the id is the identity and the text is not compared, so a changed row sent again under the same id is not stored again. Send a correction under a new id. Without row_ids, send one import_id with the row_offset each page starts at, and resend the exact bytes you sent: a number that has been re-spelled (1e2 sent back as 100.0) is a different row.

Datasets

A dataset is a saved set of conditions: tags (any or all of them), feedback, where samples came from, the model that answered, and two times. It is a live selection, not a copy. A pipeline trains on its own tags or any dataset it adds (datasets on the pipeline), narrowed by its conditions; a change applies from its next attempt. Feedback words are a person’s: a sample the judge picked is unreviewed.

Models

Pipelines

A pipeline’s id is its rule_id. versions lists only the attempts that went live, v1 first; attempts lists every attempt; live_version is the version answering now, or null. live_deployment is what the deployment the live version answers from was last seen doing, {"status", "checked_at"}: serving, starting, stopped, paused_for_funds or deleted (it answers nothing and is not billed), paused_by_platform (paused while billing could not be reached; it resumes by itself) or failed (it answers nothing). It is recorded by the platform, not asked for when you read the pipeline, so it is as fresh as checked_at; it is null when no version is live or nothing has been recorded yet. minimums says how far the data is from the next attempt, by the fixed rule. versions[] and attempts[] carry each comparison’s pass rates, candidate_pass_rate and incumbent_pass_rate. Each of attempts[] also says what was decided about it, decision (rejected when a member chose not to publish it, and expired when nobody decided before the review deadline, where its outcome says stopped; auto_rejected when the pipeline’s rule turned it down; null while nothing has been), and the samples its set trained on and held back, train_rows and holdout_rows, and its lineage, what it was built on and trained on (versions, attempts and active_run each carry one; see Building on the live version). chain is the live version as the versions merged into it, null for a pipeline that trains every version fresh. storage is what the pipeline’s history (its reports and the training rows of its sets) takes up in storage, {"bytes", "objects", "where", "billed"}, where being platform, your_storage or both, counted toward your workspace’s storage quota and deleted with the pipeline. billed is true when part of it is in the platform’s storage and is charged, as Conscious Loop storage at the storage rate with no daily minimum (each day’s cost is added up and charged once it reaches a whole cent); the part in your own storage never is. The merged models of a pipeline that builds on its live version are not part of it: they are billed as your model storage, like any trained model.

A list answers at most 5 of each pipeline’s attempts, the newest, and one pipeline at most 20, with attempts_total and versions_total counting them all: page through every attempt with GET /api/loop/pipelines/{id}/attempts. Each attempt carries explanation, {"headline", "detail", "next"}: what became of it, why, and what happens next, as sentences (next is null when nothing does). training is the pipeline’s own training settings, each null when automatic, and training_defaults the automatic values for this pipeline (min_epochs is 8 for Gemma 4, 3 otherwise): see Training settings. What each field you can send means is in Attempts and versions.

Attempts

An attempt is a training run here: attempt_no is its attempt number and version_no its version, set only if it went live. available_actions says which of approve (promote), reject, rollback and stop (cancel) the attempt allows now. A deleted pipeline’s attempts, reports and benchmark results are deleted with it: lists leave them out, and reading one answers 404.

Benchmarks

A benchmark is pinned once from 20 to 200 samples, named by id, and never edited; retire it to stop using it. From the moment they are pinned, no pipeline trains on its conversations. Attach it to a pipeline with POST /api/loop/pipelines/{id}/benchmarks.

Live Traffic & Stats

Rules in full

The guide says what each part of a pipeline does. These are the exact rules behind the held-back test, scoring and the comparison machines, word for word as the platform applies them.

Machines

You do not choose machines or prices. When you create a pipeline, change its model or switch it back on, the platform picks three to five GPU options that can train its model (cheapest first) and three to five that can run it for the comparison, from what is available. When an attempt starts and too few of them are still available at the price chosen, it picks again at that day’s prices. The challenger model gets its own. If none is free when an attempt starts, it waits in the queue, and waiting is not billed.

The new model and the live model both answer every held-back sample. While no version is live (before v1, or after a rollback left none live) the other side is the pipeline’s base model, and it answers on the new model’s own machine: that machine holds the base model’s weights with the new LoRA adapter on top, and answers once with the adapter and once without it, so no second machine is started. A challenger is compared with the pipeline’s base model, which its own machine does not run, so for a challenger the base model answers on a machine of its own beside it.

What the comparison machines cost. An attempt pays for the comparison machine’s GPU time for as long as it runs: while it answers the held-back set and the benchmarks, and also when too few held-back samples were left to compare anything. While no version is live the same machine also answers as the base model, so there is one machine to pay for. A challenger’s attempt also has the base model’s machine beside it, and so does an attempt whose machine could not answer as the base model; the latter is held to the one comparison amount rather than getting one of its own. Each machine is charged from the moment it is on a GPU to the moment it is stopped, at the price it was placed at; time waiting for a GPU is never charged, but a machine that is up waiting for the other one is. Its safety limit is 3 hours at the hourly price of the most expensive comparison option; while no version is live the same machine answers as the base model too, and a challenger’s attempt gets the same again for the base model’s machine beside it.

When an attempt is better

How many are held back. At least 100 of the pipeline’s good and corrected samples, or its share of them if that is more, rounded up. The share is 5% unless you choose more, up to 20%, in the pipeline’s Settings (or holdout.percent in the API). At 5%, 800 samples hold back 100 and 4,000 hold back 200; at 20%, 4,000 hold back 800. A larger share holds back more samples, and each training is judged on up to 1,500 of them, the ones held back longest first, so later trainings are mostly asked the same questions: a share that holds back more than 1,500 only keeps samples out of training, and they stay out even if the share is lowered later.

The new model is better on the held-back set when at least 100 samples were scored, less one for each sample the judge could not score or that was gone before the comparison started (deleted or expired while the attempt trained), to no fewer than 90; no more samples could not be scored than half the number scored; it won at least 55% of the samples where one answer was better; its average score is at least 0.05 higher, or, when that is less, a quarter of what was left to gain above what it was compared with (against a live model scoring 0.933, a quarter of the 0.067 left: +0.017), never less than the comparison can tell from chance (1.645 standard errors of the average change) and never less than 0.01; and it won more often than chance explains: a model no better than the live one would win that often less than one time in twenty. On 50 decided samples that is 32 wins or more; on 100, 59. A sample the new model never answered counts as a loss for it, not as a sample left out. With too few scored samples, too many the judge could not score, or none scored at all, the result is inconclusive, which is never put live automatically. When fewer than 90 held-back samples are left by the time the comparison starts, nothing is compared on the held-back samples and nothing is spent scoring them. A pipeline with benchmarks is still measured on them, and measuring them is charged under benchmarks. The machines already started for the comparison (the new model’s, and the base model’s when it needed its own) are still charged for the time they ran, including while they answer the benchmarks, and the attempt’s cost shows it under comparison machines. A held-back sample kept in your own storage that cannot be read at that moment is waited for, for up to an hour, and then left out and counted as missing like a deleted one; one that is gone from your storage is left out at once.

Scoring

What scoring costs. Each held-back sample takes four calls to the judge (two ratings and the comparison in both orders), one to two cents at serverless prices, paid from your wallet. Scoring one attempt may spend at most $25. When the pipeline has benchmarks, $10 of it is kept for them (each at most $5), so the held-back comparison is sized to what the rest covers: 1,500 held-back samples, which a pipeline holds back at about 30,000 good or corrected samples at the default 5%, or at 7,500 at 20% (1,500 divided by its share). A held-back set of up to 1,500 samples is scored whole. A larger one is scored on a fixed sample of exactly 1,500 of its samples, identical for every attempt, so that every version is measured on the same samples and stays comparable. If a dearer judge reaches its share first, the attempt is reported on the part of that same sample it scored, with a warning that says so.

When scoring cannot run. The calls go through an API key named Conscious Loop that the platform keeps in your workspace’s API keys. If that key has been deleted or is refused, the platform makes a new one before the next attempt starts. An attempt whose key is missing, or is refused while it is being compared, is not held open: its comparison ends at once as inconclusive, on what it had compared so far. Pressing Train now (POST /api/loop/pipelines/{id}/run) makes a new key and starts again. Nothing changes in your app meanwhile. These are the exact sentences you will see:

The attempt’s reason, key missing
What it says
the comparison could not decide: scoring could not run because this workspace's scoring key is missing (press Train now to restart it); nothing has changed in your app
The attempt’s reason, key refused
What it says
the comparison could not decide: scoring could not run because this workspace's scoring key was refused (press Train now to restart it); nothing has changed in your app
The pipeline’s next_reason, when no new key could be made before an attempt
What it says
waiting: scoring could not run: this workspace's scoring key is missing or was refused, and it could not be renewed just now. Press Train now to restart it
A benchmark on that attempt, key missing
What it says
Scoring could not run: this workspace's scoring key was missing, so the benchmark was not finished. Press Train now to restart it.
A benchmark on that attempt, key refused
What it says
Scoring could not run: this workspace's scoring key was refused, so the benchmark was not finished. Press Train now to restart it.

If Train now cannot make a new key just then, it starts nothing and answers 503 AGENT_COULD_NOT_START; press it again in a minute. Two more ends are not the key going missing. When the key reaches its monthly spending cap, nothing new is started until the next month (or until the cap is raised): the pipeline’s scoring_capped_until says when (the first moment of the next month, UTC; null when the key is not capped), its next_reason starts waiting: scoring is capped until and names the day, and Train now answers 409 SCORING_CAPPED with the same time as resumes_at. When the model the comparison is scored on is no longer offered, the comparison ends at once as inconclusive and says so, and the next attempt is scored on the model the platform chooses then.

Details by topic

The Conscious Loop guide says in plain words what a pipeline does. These are the same topics with every header, field, code and limit a developer needs. In the API a training is an attempt, and the automatic test is the held-back test.

Quick start

Start in three steps.

  1. Create a pipeline (New pipeline in the console): a name, a model or one of your deployments, and when to train. Everything else has a default, under More options. Its tag is its name in lower case with dashes, so Support replies gets the tag support-replies, and its model name is loop-support-replies. Saving it is your authorization for it to train and to charge your wallet for the GPU time it uses.
  2. Connect your app: add the header X-Loop-Pipeline: support-replies to its model calls, and send 👍, 👎 or the right answer for what they answered, or import a file of conversations you already have (Connect your app).
  3. Get v1: at 1,100 good or corrected samples the pipeline trains its first attempt and tests it. When it is better (or you approve it) it goes live as v1: point your app at loop-support-replies.
curl -X POST https://api.runbios.ai/api/loop/pipelines \
  -H "X-API-Key: $RUNBIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "Support replies", "model": {"id": "meta-llama/Llama-3.1-8B-Instruct"}}'

Recording a call

Call your model as usual, a serverless model or a deployment, on /v1/chat/completions, and add the header X-Loop-Pipeline with the pipeline’s tag. The conversation is recorded as a sample with that tag, tool calls included, and the response carries its id in X-Loop-Sample-Id (when streaming too: the header arrives before the first token). Keep the id to send feedback on it.

/v1/messages records the same way on serverless models; a deployment, loop-<tag> included, answers it 400. A tag no pipeline has (misspelt, or deleted) is not recorded under. The pipeline’s Connect tab says, in one line, which calls were not recorded and why, per tag (capture on GET /api/loop/pipelines/{id}; the whole workspace’s in GET /api/loop/stats).

curl -i https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $RUNBIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Loop-Pipeline: support-replies" \
  -d '{"model": "your-model",
       "messages": [{"role": "user", "content": "How long do refunds take?"}]}'
# x-loop-sample-id: 7c2e0b84-15da-4f37-9a6e-8d13c5f0b2a7

The header is your consent for that one call: it is recorded even when recording is off for that model. Name several pipelines with commas (X-Loop-Pipeline: support-replies, refunds). A call that also sends X-Loop-Capture: off is never recorded. If your app’s model runs somewhere else, post each conversation to the pipeline’s own source instead, POST /api/loop/traces with "deployment_id": "support-replies": creating the pipeline switched recording on for it (API reference).

  • The SDKs send the header for you: TypeScript loop: { pipeline, onSampleId }, Python loop={"pipeline": ..., "on_sample_id": ...}, on their chat and messages calls; the result’s loop_sample_id (a field, not an MCP tool) is the sample id.
  • Feedback sent right after the answer may get 409 SAMPLE_NOT_YET_RECORDED: the sample is stored a moment after the answer (after a stream ends). Retry after Retry-After seconds; the SDKs do this for you.
  • If you send X-Request-ID, keep it unique per call: a repeated one is recorded once.
  • A stream that is cut short is not recorded.

Feedback

Feedback decides what is learned, and a person’s always wins. Send it on the sample id with POST /api/loop/traces/{sample_id}/signals, or with the SDKs’ upvote, downvote and correct. In the console, open the sample on the Samples tab.

👍 Good
Body
{"verdict": "accepted"}
What is learned
The model’s own answer, as it is.
👎 Wrong
Body
{"verdict": "rejected"}
What is learned
Nothing: the answer is left out of training.
Wrong + right answer
Body
{"verdict": "edited", "correction": "…"}
What is learned
The right answer you sent.
curl -X POST https://api.runbios.ai/api/loop/traces/$SAMPLE_ID/signals \
  -H "X-API-Key: $RUNBIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"verdict": "edited", "correction": "Refunds reach your card within five working days."}'

A sample is recorded a moment after its answer: feedback that arrives first is refused 409 SAMPLE_NOT_YET_RECORDED, so retry after a couple of seconds (the SDKs do). Feedback is added, never overwritten, and the newest from a person is the one that counts. Tags can be added in the same call (labels).

Many at once: send them in parallel, or with the SDKs’ signalMany / signal_many (example). Sending a sample yourself? Its feedback can ride along (feedback on POST /api/loop/traces).

Importing a file

Upload JSONL, one conversation per line, on the Connect tab, or send the rows to POST /api/loop/pipelines/{id}/import, up to 5,000 per call (the console sends a bigger file in pages). Each line is messages, and everything in it is kept: system and developer instructions, the user’s turns, and tool calls with their results. The last assistant turn is the answer it learns. Every row gets the pipeline’s tag. Choose “these answers are good examples” ("default_feedback": "good") to learn from the whole file; otherwise its rows wait for review.

json
{"messages": [{"role": "system", "content": "You answer for our support desk."}, {"role": "user", "content": "How long do refunds take?"}, {"role": "assistant", "content": "Refunds reach your card within five working days."}]}
{"messages": [{"role": "user", "content": "Where is order 48812?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "find_order", "arguments": "{\"order_id\": \"48812\"}"}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "{\"status\": \"shipped\"}"}, {"role": "assistant", "content": "Order 48812 left the warehouse this morning."}]}

Other row shapes (a prompt and a completion, a correction, a chosen and a rejected answer, a question with its right answer), retrying an upload that failed halfway, and what the answer counts are in the API reference.

Recording every call of a model

Advanced: instead of the header, switch recording on for a deployment or serverless model with the pipeline’s tag, on the Connect tab or with PUT /api/loop/configs/{deployment_id} (enabled, tags, and optionally sample_rate, the share of conversations kept, and retention_days). Every conversation it answers is then recorded with that tag. It still needs feedback before it teaches anything. When a new version goes live on a new deployment, the recording setting moves with it.

Calling the improved model

Once v1 is live, point your app at the pipeline’s model name, loop- and its tag. It always answers with the live version, so your app keeps the same URL and model name from one version to the next; the name never changes, even when you rename the pipeline. The API key needs the deployments:read scope. Keep the header to go on collecting samples.

bash
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $RUNBIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Loop-Pipeline: support-replies" \
  -d '{"model": "loop-support-replies",
       "messages": [{"role": "user", "content": "How long do refunds take?"}]}'

Before v1, and while the live version is not serving, the name answers nothing (see Serving and billing).

Samples, tags and scoring

A sample is one conversation your app had, plus your feedback. A pipeline trains on every sample that carries any of its tags (or is in one of its datasets), and only on the ones marked Good or Wrong + right answer, or picked by automatic scoring. One marked Wrong is kept and shown but teaches nothing.

  • The Samples tab counts them (total, Good, Wrong, Wrong + right answer, No feedback and held back; in a pipeline that builds on its live version, also learned and not yet learned) and lists them 25 at a time, filtered by All, Good, Wrong, Wrong + right answer or No feedback, by tag when the pipeline has more than one, or by a search of their words (2 to 200 characters, over every message, the answer and its tool calls). A sample kept in your own storage cannot be searched, and the tab says how many a search left out. A chart shows how many arrived each day over the last 30 days, by feedback. Open one to read the whole conversation, give feedback and change its tags; it also says which pipelines it feeds.
  • Tags (labels in the API) are lower case: letters, digits, ., _ and -, up to 63 characters, at most 32 on a sample. A pipeline trains on its own tag and on any more tags, or datasets, listed in its Settings.
  • One sample, several pipelines: a sample feeds every pipeline that has any of its tags. Two pipelines can share a tag, but no two have the same own tag.
  • A correction to a sample a version already learned, in a pipeline that builds on its live version, is new work: the next attempt learns the corrected answer on top of what is live. The old answer is not unlearned; it fades as the correction is learned, and is gone only when the model is rebuilt from the base.
  • Deleting a sample takes it out of every later attempt, and can rebuild a model that learned it (Deleting a sample a version learned). Samples expire by their source’s retention setting (Privacy).

When a pipeline does not yet have the samples its next attempt needs, the judge that tests its attempts also scores samples nobody has reviewed, newest first, with the same rubric. A sample it scores 0.8 or more is trained on like one marked good; one below is not. Your feedback always decides: a sample a person reviewed is never scored, and reviewing one it scored replaces its score. Scored samples stay under No feedback on the Samples tab: a pick is not feedback, andPicked by the judge, Below the bar and Not scored yet split them. A pipeline whose feedback condition leaves No feedback out is not scored. Scoring stops as soon as the next attempt has what it needs, and what it cost shows on that attempt as Scoring unreviewed samples. It is on by default; the switch is in Settings, where it runs (auto_scoring on the pipeline in the API).

What each kind of feedback trains. Every attempt trains with SFT (supervised fine-tuning on examples) today: the other methods are switched off on the training engine and nothing trains with them. Feedback is kept as long as its sample is, and samples expire by their source’s retention setting.

👍 Good
Trains today (SFT)
The answer, as it is
Methods it suits, not offered yet
KTO
👎 Wrong
Trains today (SFT)
Nothing: left out
Methods it suits, not offered yet
KTO
Wrong + right answer
Trains today (SFT)
The right answer you sent
Methods it suits, not offered yet
DPO
Gold (gold in the API)
Trains today (SFT)
The right answer you sent
Methods it suits, not offered yet
DPO, GRPO
No feedback
Trains today (SFT)
Nothing, unless automatic scoring picks it: then the answer, as it is
Methods it suits, not offered yet
None

Attempts and versions

The Conscious Loop turns what your model actually did into what it learns next. It records the samples your model answers, learns from the ones you mark good or correct, and trains a new version of the model on its own. A new version goes live when it beats the live model on samples it never trained on, or when you approve it, or, in a pipeline that builds on its live version, after a deleted conversation, when the rebuild without it is not clearly worse (unless you chose Only when I approve). Version numbers count up: v2 went live after v1.

Every time a pipeline trains, that training is numbered #1, #2 and so on, and tested (Tests). A training that goes live becomes the next version; one that does not changes nothing in your app. Version numbers count up and never change.

When it trains. The first attempt waits for 1,100 good or corrected samples: 1,000 to train on and 100 held back to test it (1,250 if the pipeline holds back 20%). Each next attempt waits for 1,000 new ones since the last attempt. These minimums are fixed for every pipeline, so that each attempt has enough to be worth testing, and the pipeline says how many are still missing. A schedule (daily or weekly, at an hour you choose) is optional: with one, an attempt starts at the next scheduled time once the minimum is met. Train now skips the schedule, never the minimum. A pipeline runs one attempt at a time.

What it trains. Every attempt fine-tunes a LoRA adapter on the pipeline’s good and corrected samples, except the held-back ones: on the pipeline’s base model, or, in a pipeline that builds on its live version, on top of the live version. A pipeline that builds on its live version trains each next attempt as a new LoRA adapter on top of the live version, on the samples it has not learned yet plus a replay of what it already learned, picked at random to keep it, and merges a winner in (Building on the live version). Any other pipeline trains each version fresh, on all of its good and corrected samples so far (the newest 200,000). The platform picks the GPUs, and how each attempt trains unless you set it (Training settings).

Each attempt says why it was made: enough new samples (data), the schedule, Train now (manual), the second model (challenger), or rebuild (a conversation the live version learned from was deleted: see below).

The model is one from the list a pipeline offers: every model the platform can train with LoRA and serve, recommended first. Settings → Model also lists the models not offered here, each with the reason. The exact revision is pinned when you save, and the model cannot change while an attempt runs. After a change the numbering carries on: the first attempt on the new model to go live is the next version.

Version limit. By default a pipeline keeps making versions for as long as they win. A limit from 1 to 100 stops it after that many versions (attempts that do not go live do not count), and the pipeline is then complete: raise or remove the limit to carry on.

Building on the live version

A pipeline can build each version on the one before it. Once a version is live, each next attempt waits for 1,000 new samples and trains a new LoRA adapter on top of the live version, on the samples it has not learned yet: the new ones, and those of attempts that did not go live. So that it does not forget, it also replays a random share of what the live version already learned: half of what the attempt trains on by default (as many as it has not learned, or all the live version learned when that is fewer), less if you lower Replay of learned samples in Training settings. When an attempt wins, or you approve it, its adapter is merged in and it goes live as one model: the pinned base with every version merged into it, oldest first.

  • Each version says what it is made of, for example v3 = v1 + v2 + its own training on 1,000 new samples (chain on the pipeline and lineage on each attempt in the API).
  • The first version answers as an adapter on the base model; from the second version on, the live version is one merged model. A model whose adapters cannot be merged is not built on: each attempt of it trains from the base on all samples, and says so.
  • A merged model is a full copy of the model. It is kept while it is live and while a rollback can still put it back, and deleted with the rest of an attempt that did not go live, within a day. It is billed as your model storage, like any trained model, for as long as it is kept. Each version’s own adapter is billed too, as checkpoint storage, like any training checkpoint. The pipeline’s reports, benchmark results and training rows are billed only for the part kept in the platform’s storage, as Conscious Loop storage, where the pipeline’s storage says billed.

Rolling back

For 30 days after a version goes live you can roll it back with Go back to v4 on its card on the Versions tab, or on its own page: your model name returns to the version that was live before it. When none was (rolling back v1, say), the name answers nothing until another version goes live, and the next attempt is compared with the pipeline’s base model as it is then. A rollback never changes the models, so after a second model’s version went live that is the second model’s base; the rollback dialog names it. Only the most recent promotion can be rolled back, and numbers do not change.

In a pipeline that builds on its live version, the version put back is what the next attempt builds on, and the next attempt trains again on what the rolled-back version learned. That holds when the version put back is on the pipeline’s model now; one on the model before a second model took over cannot be built on, so the next attempt trains the pipeline’s model from its base on every sample.

  • A version cannot be rolled back while an attempt is being built on it or compared with it (409 RUN_ACTIVE). The console offers Cancel attempt N and roll back, which stops the attempt first: you pay only for the GPU time it used, and its samples go to the next attempt.
  • A version that learned from a conversation you deleted, or was built on one that did, cannot be put back (409 ROLLBACK_DELETED_DATA).

Deleting a sample a version learned

In a pipeline that builds on its live version, deleting a conversation the live version learned from rebuilds the model without it: the next attempt trains from the base on every remaining sample, without waiting for new ones (deletions close together are rebuilt once), and Train now starts it at once.

  • It needs 1,000 samples to train on; with fewer it waits and says how many remain. A paused pipeline rebuilds when it is switched back on, and the version limit does not hold a rebuild.
  • It is compared with the live version on the same held-back samples, and goes live unless it is clearly worse; if it is, a member is asked. A test that could score nothing showed nothing worse, so the rebuild goes live, unless it gave no answers itself. Under Only when I approve it waits for your approval like any attempt, whatever it scored. Until a rebuild goes live, the live version still holds what it learned from the conversation.
  • Once a rebuild is live, no version that learned the conversation can be rolled back to, and their stored weights are deleted.
  • A conversation that expires under its source’s retention setting rebuilds nothing: what a version learned from it stays.

Second model

A pipeline can also try a second model (Settings → Second model, challenger_model in the API), a different model from its own. After each training of the pipeline’s model, the second model trains on the same samples and is tested the same way, against whatever is live then. Each model keeps its own line of versions: the first time, the second model trains from its original model on all samples; after that, where the pipeline builds on its live version, it builds on its own latest version with the samples new since that version.

Each training that does better goes live in turn, and it is billed like any training. When it goes live, it takes the pipeline’s place: it becomes the pipeline’s model and the one it replaced becomes the second model. Remove it with "challenger_model": null; a training of it already running still finishes. The pipeline’s second_model says where it stands, and why when it did not train on a set.

Training settings

Automatic (recommended) lets the platform decide how each attempt trains, from how many samples it has and how long they are. Custom sets any of the five values below yourself and leaves the rest automatic. A change applies to the next attempt: earlier attempts keep the settings they trained with, and each attempt’s report shows them, marked Automatic settings or Custom settings.

Passes over the data (epochs)
Automatic
Enough for about 1,100 training steps of about 16 samples each: at least 3 passes and at most 15, and at least 8 for Gemma 4
Custom
1 to 15
Learning rate
Automatic
0.0002, easing off towards the end
Custom
0.000001 to 0.01
LoRA rank
Automatic
16
Custom
4, 8, 16, 32 or 64
LoRA alpha
Automatic
32
Custom
1 to 512
Replay of learned samples
Automatic
50%: half of what an attempt that builds on the live version trains on
Custom
0 to 50%

Less replay trains faster and forgets more of what earlier versions learned. As it trains, an attempt is checked about ten times on 5% of its rows kept aside (never the held-back samples), stops early after three checks without improving, and keeps the checkpoint that did best (its last save is never checked, so it is used only when no checked checkpoint was kept).

Gemma 4 trains differently. It gets at least 8 passes, never stops early, and keeps its last checkpoint. At 3 passes it learned the form of the answers but few of the facts in them. Its check score also gets worse while it is still learning facts, so the checkpoint that “did best” on the checks knew fewer of them than its last one (10% against 53% of the facts in one test).

In the API this is training on the pipeline: null for a value means automatic, "training": null makes them all automatic, and leaving it out keeps what is set. training_defaults gives this pipeline’s automatic values, and a value out of range is refused with 422 TRAINING_SETTING_OUT_OF_RANGE, naming it.

json
"training": { "epochs": 8, "learning_rate": 0.0002, "lora_rank": 16, "lora_alpha": 32, "replay_percent": 50 }

Tests and going live

Every training is tested before it can go live: always on the automatic test, which needs nothing from you, and on each of your benchmarks, if you add any. The Tests tab has one row per test: its newest result, a small chart of the newest trainings (on the automatic test each beside the version it was tested against), and See questions. Going live turns the results into one decision.

An answer is right when the judge rates it at least 7 out of 10 against the right answer, or on its own when the question has none (on an exact-match benchmark, when it matches). Not measured means there is no result, never a zero: the benchmark was added after that training, or its test did not finish, and the report says why. Each training is tested on a temporary GPU started and stopped automatically, which answers conversations of up to 32,768 tokens (or the model’s own limit): a longer one counts as a loss for the training.

Questions from your own samples that the model never trained on. Before a training can go live, it answers them, and so does the live version (before any version is live, the pipeline’s original model). The judge compares the two answers with the right answer, and the training counts as better only when it wins clearly (the exact rules).

  • Which samples: 5% of the pipeline’s good and corrected samples, or up to 20% (Settings → Advanced → Test share), and at least 100. A sample kept aside is never trained on and stays aside for every later training, so the set only grows. Each training is tested on up to 1,500 of them, so once more are kept aside every training answers 1,500. The set grows with the pipeline, so two trainings are seldom tested on exactly the same samples: the one score that follows every version is a benchmark’s. New samples fill the kept-aside set first, so raising the share delays the next training.
  • How it is checked: the judge rates each answer against the right answer a person gave the sample (the answer they marked good, or their correction; feedback sent through the API or an SDK is always a person’s, so an upvote or a correction sent that way is one) and on the rubric alone when it has none. The two answers are also compared side by side, twice with their order swapped, and the judge is never told which is live: a sample is a win only when both readings pick it.
  • Better means all of these: at least 100 samples scored (down to 90 when some were deleted or could not be scored), no more failures than half of those, wins on at least 55% of the samples where one answer was better, an average score at least 0.05 higher (or, once what it is compared with scores above 0.80, a quarter of what is left to gain: +0.017 against a live model at 0.933), and more wins than chance explains (less than one time in twenty). Otherwise it is not better, or inconclusive when too little could be scored, and an inconclusive training never goes live on its own.
  • The judge is chosen when you create the pipeline, from the chat models your workspace can call (anthropic/claude-sonnet-5 whenever it is offered). Its calls go through an API key named Conscious Loop kept in your workspace, and cost one to two cents per sample, at most $25 a training. If that key is missing or refused, the test ends inconclusive and says so; Train now makes a new key and starts again.

A benchmark is a fixed set of 20 to 200 conversations you choose, such as questions with known answers. From the moment you add it, every training is tested on the same set, and so is the live version, so you can see whether each version gets better. Add one on the Tests tab, from a JSON or JSONL file or from samples you pick. The first in the list is your main benchmark, which the progress chart follows.

  • Scored by the judge (the default): each answer is rated against the right answer, or on its own when the question has none, as in the automatic test. Exact match: the answer must be the right answer, ignoring case and spacing, or the same number (12.00 and 12), or make the same tool calls with the same arguments; no model is asked, and every conversation needs a right answer. It suits a number, a label or a tool call: written answers rarely match word for word, and a value hidden when the conversation was recorded ([redacted:email] and the like) cannot be matched.
  • The right answer is the newest a person gave the conversation: the answer marked good or the correction (in a file, ground_truth or correction). The benchmark keeps its own copy, so later feedback does not change it, and no pipeline trains on its conversations from the moment they are pinned: pin conversations no training has learned yet.
  • A score is the average over the conversations scored, published only when no more than 1 in 20 could not be scored. A benchmark improved when the new score is higher and more conversations scored higher than chance explains. Older versions are not measured again on a benchmark added later: they show not measured.
  • A benchmark cannot be edited: make a new one. Taking it off a pipeline, or retiring it, keeps every score it produced.

The pipeline’s going-live choice (on the Tests tab, Change beside its one line; promotion in the API) turns the test results into one decision. Under every choice but manual, a training that is better goes live on its own, one that is not better is closed, and an inconclusive one waits for you.

holdout
A training is better when
The default: the automatic test says so.
all
A training is better when
The automatic test and every benchmark improve.
primary
A training is better when
Your main benchmark (the first in the list) improves and the automatic test does not get worse. Needs a benchmark.
k_of_n with k
A training is better when
At least k of the tests (the automatic test and each benchmark) improve.
weighted
A training is better when
The average change, weighted 1 for the automatic test and 1/n for the nth benchmark in the list, is above 0.05, and at least one test improved.
manual (Only when I approve)
A training is better when
Nothing goes live on its own: every tested training waits for you, a rebuild after a deleted conversation included.

One training is decided differently under every choice but manual: a rebuild after a deleted conversation goes live unless it is clearly worse than the live version (above). Under manual it waits for you like any training.

A training waiting for you is approved (it becomes the next version) or rejected on its own page (open it from the Versions tab); approving an inconclusive one asks you to confirm ("force": true in the API). While it waits its test machine is stopped, so waiting costs nothing, and after 14 days without a decision it is closed. Going live points the pipeline’s model name at the new version.

Every status says what happened, why, and what happens next, for example Could not be tested: the comparison machine did not finish starting, so the attempt was stopped and nothing changed. The next attempt tries again. (explanation on each attempt in the API).

running
Meaning
Training, or being tested.
became_version
Meaning
It went live as the version shown, because it won, because you approved it, or, after a deleted conversation, because the rebuild without it was not clearly worse (unless you chose Only when I approve).
not_better
Meaning
Tested, and not better than the live version.
inconclusive
Meaning
Tested, but the test could not decide: too few samples could be scored, or too many failed. Never put live on its own.
not_tested
Meaning
It ended before it could be tested: its comparison machine never started, or it expired first. Nothing changed.
waiting_for_review
Meaning
Tested, and waiting for you to approve or reject it.
stopped
Meaning
Stopped by you or by a limit, or rejected after it was tested.
failed
Meaning
Something went wrong; the attempt says what, and what happens next.

Compare puts any two attempts of a pipeline side by side, live or not, or any two versions. Held-back numbers are shown side by side and never subtracted: each attempt was tested against what was live on its own day. Benchmarks both were measured on show both scores and the change. The same held-back conversations lists the conversations both were scored on, each with both scores and whether the second scored higher, lower or about the same (0.02 or more above counts as higher); each attempt’s pass rate over these is the like-for-like number. Two attempts scored by different judges are shown without saying which is higher.

Costs and limits

The live version runs on its own deployment: the machine that answers when your app calls the pipeline’s model name. Serving is whether that machine is on. While it is on it is billed by the hour, like any deployment; turn it off or on from its deployment (linked on the Overview). While it is off, your app’s calls to the model name fail, nothing is billed, and Train now is refused with 409 LIVE_VERSION_STOPPED, because an attempt could not be tested against it. When a newer version replaces it, the old machine is stopped ten minutes after the switch; a rollback stops it once your model name no longer points at it.

An attempt costs what it used, paid from your wallet; you do not choose machines or prices, and waiting for a GPU is never billed:

  • training: its GPU time, as metered;
  • the comparison machine: the GPU time it served, also while it answers your benchmarks, which is what your statement charges. Until billing has settled the machine, shortly after it stops, the attempt counts it from the moment it is on a GPU, a few minutes before your statement starts billing it, and shows that figure as up to; billing’s figure then replaces it, and the attempt’s timeline says what the machine was charged. A version keeps up to: its own machine is the one that went live. A comparison machine that never started costs you nothing: any GPU time metered for it is returned on your statement. An attempt that needs a second machine (a second model tested while no version is live) can bill the one that started for up to 20 minutes per try while it waits for the other;
  • scoring: the judge’s calls.

Each attempt shows its GPU time and cost. Merged models and adapters are billed as storage (see Building on the live version). Safety limits stop a runaway attempt, and an ordinary one never comes near them: 4 hours of training and 3 of comparison machine at the dearest GPU option chosen for it, $25 of scoring (of which $10 is kept for benchmarks), and, per pipeline per calendar month (UTC), 30 times the most one attempt may cost: a pipeline that would pass that waits until the 1st.

An attempt starts only once the wallet that pays for it holds enough for its whole training, beside your other pipelines’ trainings still running, because training is never stopped part way for money. Until then it shows Not started: waiting for funds (needs_funds) with nothing booked and no amount shown, starts by itself once the wallet holds enough or automatic top-up is on, and after 24 hours ends (INSUFFICIENT_FUNDS:NOT_STARTED) and pauses the pipeline. If your wallet runs out later in an attempt, the attempt waits up to 24 hours for a top-up (needs_funds), then stops and the pipeline pauses: add funds and press Switch back on in Settings. When no GPU is free the attempt shows Waiting for GPUs and books again after 30 minutes and 1, 2 and 4 hours, ending after four tries or 24 hours (COMPARISON_NO_GPUS). A comparison machine that gets no GPU at all says since when it waits and when the wait ends, and after a day the attempt stops (CANDIDATE_NOT_STARTED:NO_GPU): that machine never started and is not charged. When your workspace is at its limit of running deployments it shows Waiting for room: stop a deployment and press Train now.

Pausing and deleting a pipeline

Pause in Settings ("enabled": false): nothing trains until you switch it back on, and whoever switches it back on is the member it then runs on.

Deleting a pipeline stops an attempt in progress, switches recording off for the source named after its tag, and stops its live version: your model name stops answering, and the deployment it answered from is stopped and deleted, so it is no longer billed, unless another pipeline or model name of yours answers from that deployment too. If one does, that deployment keeps running and billed until you stop it on Deployments, and its trained checkpoint is kept (and billed as storage) until that deployment is deleted. Its versions, attempts, reports and benchmark results leave the console and every list at once, and reading one answers 404. Then, in the background, the text of every conversation in its reports is erased, and its trained checkpoints (including any an attempt saved as it was stopped), the rows it trained on and everything it kept in storage are deleted. The exception is the checkpoint of a deployment left running because another pipeline or model name of yours answers from it too: that checkpoint is kept (and billed as storage) until that deployment is deleted. The platform keeps only numbers for its own accounting, such as what each attempt cost, and the loop shows none of them. Its samples stay, and its name stays reserved.

Privacy and storage

What the loop records, what it removes first, who can see it, and how long it is kept.

  • Nothing is recorded without your say-so. A call is recorded when it carries X-Loop-Pipeline (your consent for that one call), or when recording is switched on for its source: by you, or by the member who creates a pipeline, for the source named after its tag. The setting records who switched it on and when. A call with X-Loop-Capture: off is never recorded. Switching recording off, or deleting the pipeline, stops recording and keeps what is stored.
  • Secrets are removed before anything is stored. Email addresses, API keys and tokens, AWS keys, JWTs, private keys, card numbers, US social security numbers and IP addresses are replaced with a marker such as [redacted:email]: in the conversation, the answer, tool-call arguments and corrections.
  • Recorded samples expire after the source’s retention period (30 days unless you set 1 to 3,650) and are deleted. Imported samples are kept for 365 days. An expired sample leaves every set, so no attempt trains on it again, but what was already made from it keeps it: its row stays in the stored copy of a set already built until that set is deleted (and in the copy an attempt registered for training), and a report that measured it keeps it until its pipeline is deleted, as a benchmark that pinned it keeps its copy.
  • Deleting a sample removes it and its feedback, and no later attempt trains on it. Versions already trained are not changed, except that in a pipeline that builds on its live version, deleting a sample the live version learned from rebuilds that pipeline’s model without it (Deleting a sample a version learned). A sample deleted, by you or by expiry, leaves every set, so no attempt trains on it again; its row stays in the stored copy of a set already built until that set is deleted with its pipeline (or, for a set no single pipeline claims, such as one built before sets recorded their pipeline, with the workspace), as it does in the copy an attempt registered for training.
  • Only your workspace can see its samples. A sample in another workspace answers as if it did not exist.
  • Where reports are kept. Each conversation of a report (the prompt, both answers, how each was scored and the side-by-side verdict), each of a benchmark’s pinned conversations, and the training rows each attempt is built from are kept in storage, in one place per organization, workspace and pipeline (a benchmark’s pinned conversations in one the workspace’s benchmarks share); the loop’s database keeps only the numbers and an index. The one exception is an older report or benchmark holding a single conversation too large to keep in storage: that report stays in the loop’s database and is erased with its pipeline, and that benchmark’s pinned conversations stay there and are erased with the workspace. What a pipeline’s history (its reports and its training rows) and a benchmark’s pinned set take up is shown on each (storage), and it counts toward your workspace’s storage quota. What is billed: only what a pipeline keeps in the platform’s storage (its reports and its training rows), as Conscious Loop storage at the storage rate with no daily minimum: each day’s cost is added up and charged once it reaches a whole cent. It is charged only where the pipeline’s storage.billed is true; an environment can measure it for a while before it charges it, and says billed false meanwhile. Deleting the pipeline stops the charge. A benchmark’s pinned conversations are never billed. What you keep in your own storage carries no storage charge from us, and the index the loop keeps in its database is never billed. The copy of a set’s training rows that each attempt hands to its training job is not billed as a dataset: those rows are billed once, where the set is kept. The copy still counts toward your storage quota until its pipeline is deleted. A pipeline’s is deleted with the pipeline; a benchmark’s pinned set is the workspace’s, and is deleted with the workspace, not when the benchmark is taken off a pipeline or retired. If storage cannot keep something just then, the attempt waits and says so.
  • Deleting a pipeline deletes its history (Pausing and deleting). Deleting a workspace does the same for every pipeline in it, and erases its benchmarks’ pinned conversations. Pipelines deleted earlier, before deleting did all of this, are cleaned up the same way. The samples themselves stay until they expire or you delete them.
  • Tests send held-back and benchmark conversations to your models and to the judge through Run BiOS serverless inference, on your workspace’s Conscious Loop key. Those calls are never recorded as new samples.

A workspace can keep what the loop stores in a bucket or container it owns: connect one under Integrations → Your own storage (Amazon S3, Azure Blob Storage or Google Cloud Storage), run its connection check, and make it the default (how). From then on the conversations the loop records, and the stored detail of its reports, are written there under your folder; the loop’s database keeps the index and your feedback (ids, tags, verdicts and corrections, scores).

  • Each pipeline’s storage.where says where its history is: platform, your_storage or both. What is in your storage is not billed as ours.
  • A sample kept there cannot be searched by its words, and while your storage cannot be read the sample says so instead of showing its text.
  • Deleting a sample or a pipeline deletes what the loop wrote there, one object at a time from its own record of what it wrote; it never lists or deletes anything else in your storage.

SDKs

Both SDKs cover the whole flow with the same names, snake_case in Python and camelCase in TypeScript. API keys need the loop:read scope to read and loop:write to change anything; neither is in the Read Only preset, because what they read is your raw conversations. To record a call your app makes, pass loop to the SDKs’ chat and messages calls, or send the header from any OpenAI- or Anthropic-compatible client, as in Record with one header.

from bios import RunBiOS

client = RunBiOS()  # RUNBIOS_API_KEY

# A pipeline, with two training settings of your own (the rest automatic)
model = client.loop.list_models()["models"][0]["id"]
pipeline = client.loop.create_pipeline(
    name="Support replies",
    model={"id": model},
    training={"epochs": 8, "lora_rank": 32},
)["pipeline"]
pid = pipeline["rule_id"]

# Feedback on a sample your app recorded with X-Loop-Pipeline
client.loop.correct(sample_id, "The answer it should have given.")

# Its attempts, newest first, 25 at a time
page = client.loop.list_pipeline_attempts(pid, limit=25, result="became_version")
for attempt in page["attempts"]:
    print(attempt["attempt"], attempt["version"], attempt["outcome"])

Beyond the calls above, both SDKs read which pipelines a sample feeds (get_trace_pipelines / getTracePipelines) and the attempt-by-attempt history behind the charts (get_metrics_history / getMetricsHistory). list_evaluation_items / listEvaluationItems pages through the held-back samples behind an attempt’s report and list_benchmark_run_items / listBenchmarkRunItems through the conversations of one benchmark measurement, each filtered by which answer won, which side passed, whether it had a right answer and whether it was scored (and, for the held-back samples, whether the two side-by-side readings agreed), sorted by the change in score, or searched with q, which says in searched how many rows it read. compare_attempts / compareAttempts puts any two attempts of a pipeline side by side, put live or not, as compare_versions / compareVersions does for its versions.

MCP tools

From an AI assistant, the MCP server has a tool for every route in the API. Saving a pipeline and training now spend from your wallet, so the assistant asks you before it calls them.

Models
Tools
loop_list_models
Pipelines
Tools
loop_list_pipelines, loop_get_pipeline, loop_create_pipeline, loop_update_pipeline (with training), loop_delete_pipeline, loop_run_pipeline (train now), loop_import_into_pipeline, loop_list_pipeline_attempts, loop_compare_versions, loop_compare_attempts
Samples
Tools
loop_chat (ask a model and save the conversation, with its sample ID), loop_capture_trace (with labels as its tags), loop_import_rows, loop_list_traces, loop_get_trace, loop_delete_trace, loop_get_trace_pipelines
Datasets
Tools
loop_list_datasets, loop_get_dataset, loop_preview_dataset, loop_create_dataset, loop_update_dataset, loop_delete_dataset, loop_export_dataset
Feedback and tags
Tools
loop_add_signal, loop_list_signals, loop_label_trace, loop_remove_label, loop_list_labels
Attempts
Tools
loop_list_training_runs, loop_get_training_run, loop_get_training_curve, loop_promote_training_run, loop_reject_training_run, loop_rollback_training_run, loop_cancel_training_run, loop_retest_training_run, loop_get_evaluation, loop_list_evaluation_items
Benchmarks
Tools
loop_create_benchmark, loop_list_benchmarks, loop_get_benchmark, loop_list_benchmark_items, loop_get_benchmark_history, loop_get_benchmark_run, loop_list_benchmark_run_items, loop_benchmark_questions, loop_retire_benchmark, loop_list_pipeline_benchmarks, loop_attach_pipeline_benchmark, loop_move_pipeline_benchmark, loop_remove_pipeline_benchmark
Live traffic and stats
Tools
loop_set_capture (with tags), loop_list_capture_settings, loop_stats, loop_get_metrics_history

For example, “mark that answer good” is loop_add_signal with verdict: "accepted" on the sample id your app got back. See the MCP Server page to connect it. The loop tools are registered only where the loop is offered: the server run with npx runbios-mcp asks the API it points at when it starts and lists them wherever the API offers the loop to its key (point RUNBIOS_BASE_URL at https://api.runbios.ai; a key without the loop:read scope gets none), and RUNBIOS_LOOP_COMING_SOON set to false or true overrides what it asks. Where the platform has the loop switched off, npx runbios-mcp still lists the five tools you need to stop paying: loop_list_pipelines, loop_get_pipeline, loop_update_pipeline (switching a pipeline off only), loop_delete_pipeline and loop_cancel_training_run. The hosted connector named there lists the loop tools.

Run BiOS Documentation. Need help? Email contact@runbios.ai