Galaris

The Lab: compare on your work cases

Capture real situations, compare candidates and check scores against human review. Each campaign retains its context and results.

Considering a different model or harness? The Lab helps you measure differences on your own work cases. Compare models and mechanism configurations in comparable contexts, then check the selected engine in use. These evaluations do not replace a full trial of tools and external effects. Choose models and harnesses.

Eleven mechanisms to examine separately

Family Mechanisms
Prepare work Dispatcher, Briefing, Planner
Organise and learn Topic detection, memory extraction, learning
Follow progress Goal follow-up
Execute Task, conversation and voice executors
Diagnose Task analysis

The incident journal is separate. Briefing remains testable and visible in history; its presence in the Lab does not establish its use in every current workflow.

Laboratory mechanisms — Test benches for the dispatcher, planner, memory and executors.
Laboratory mechanisms — Test benches for the dispatcher, planner, memory and executors.C43

Build a dataset around a question

Create mechanism-specific cases with a tested variable, context and expected result. Capture a Task, round or voice turn with provenance. Differences between source configuration and dataset parameters are visible before import.

Separate work, validation and holdout sets. Categorise nominal cases, ambiguity, missing context, multilingual input, robustness, real incidents or valid alternatives. Incomplete cases remain drafts. Model-proposed references need review.

Execute, then judge

Select candidate and judge separately. Launching fixes cases, parameters and models. The first pass produces outputs and objective checks; the second judges them. The candidate never receives the expected result.

Inspect scores by dimension, explanations, critical failures and unjudged cases. A finished campaign is not necessarily successful; a judge failure leaves a missing score.

Measure stability and resume work

Repeat each item from 1 to 20 times to compare success rate, average, extremes and spread. Cancel while preserving published results, resume remaining items with the snapshot, or rejudge saved outputs without calling the candidate.

An optional budget covers candidate and judges when admitting new steps. Repetition measures stability on these cases, not every future situation.

Add independent human review

Review can hide model identity and automated scores. Each reviewer scores and explains their decision before revealing the judgment. Disagreements become visible. A revealed review is fixed; rejudging opens another campaign.

Human review in the laboratory — A saved output is presented for blind human assessment; no review was submitted.
Human review in the laboratory — A saved output is presented for blind human assessment; no review was submitted.C52

Turn an incident into a check

Analyse a Task dossier with human context without replaying execution. Capture the diagnosis to evaluate that mechanism too.

Simulated tools do not prove external effects. The voice Lab evaluates a transcript, not STT/TTS acoustic quality. Incident tracking complements this evaluation.

Evaluate genuinely permitted choices

The Dispatcher Lab examines the modes allowed by the harness and retains deterministic decisions and their reasons. Briefing remains testable here while disabled in current Task policies. The Lab’s eleven evaluable mechanisms are separate from Dream’s twelve background mechanisms.