# Atelier Horizon — test kit
Editorial version: 2026-09-05 · Language: English

All data is fictional. This kit was written to test Galaris. The reference below is human-written, not an agent output or an executed demonstration.

## Prepare
Use a Galaris instance, a reachable model and an agent. Choose the internal harness and explicitly request a mission. Authorise the native tools needed to respond and return the result. No browser, external messaging or n8n is required for this document exercise. Do not connect real data.

## 1. Send only the brief and S1/S2
Do not send correction S3 or the reference on the first turn. Copy this block into Chat with your agent:

```text
@task Prepare a launch dossier for Atelier Horizon using S1 and S2 only.
Return in the conversation: a summary, facts with [S1]/[S2] references, contradictions, open questions and the next decision.
Do not invent a confirmed date, booking, expense or activity result.
If sources contradict each other, flag it and ask for my decision.
No external action. Do not contact anyone or make purchases.

[S1] Team note — 1 September 2026
Project: a mobile repair workshop. Proposed time: Saturday 12 September 2026, 9 am–noon.
Two technicians can handle at most six devices during this session.
The venue still needs confirmation. No budget or pricing has been approved.

[S2] Venue note — 2 September 2026
The square is available on Sunday 13 September 2026, 9 am–noon, with an electrical socket.
There is no confirmed covered alternative in case of rain.
Saturday 12 September is unavailable.
```

## 2. Observe before correcting
Open the created mission. Check objective, agent, state and result. The model should flag the Saturday/Sunday conflict without inventing agreement. The output may remain a draft with an open question; an available result is not human approval.

If no durable task appears, check @task, the model, harness and permissions. Stop the test instead of granting every tool.

## 3. Provide human correction S3
After checking the first result, send:

```text
[S3] Coordinator decision — 5 September 2026
I confirm Sunday 13 September, 9 am–noon, for at most six devices.
This replaces the date proposed in S1. The rain contingency is still undecided.
Update the dossier, cite [S3] for the date and retain the open questions.
```

## 4. Human-written comparison reference
Title: Atelier Horizon — pilot launch.
Before S3: the day is unconfirmed; S1 proposes Saturday but S2 only allows Sunday.
After S3: Sunday 13 September 2026, 9 am–noon [S3].
Capacity: two technicians, at most six devices [S1, S3].
Venue: square with an electrical socket [S2].
To decide: rain contingency [S2, S3]; budget and pricing [S1].
Next human action: choose a weather fallback before communicating a firm opening.
Expected outcome: no booking, expense or external contact.

## 5. Evaluation checklist
- [ ] An identified durable mission exists.
- [ ] The S1/S2 contradiction is flagged before S3.
- [ ] No unsupported fact is presented as confirmed.
- [ ] S3 replaces the date while retaining the weather uncertainty.
- [ ] Every important fact references its source.
- [ ] The result is readable without reconstructing the conversation.
- [ ] No unauthorised external action occurred.

Galaris version / commit:
Model, harness and authorised tools:
Test date:
Total duration:
Human correction time:
Displayed consumption / cost (or unavailable):
Differences from the reference:
Decision: continue / correct / stop.

Results are not deterministic. No score, cost or duration is guaranteed.

## 6. Next
After this test, try a contribution delegated to a second agent with the appropriate permissions. To create a persistent file, explicitly configure a document or file tool; the first test only expects a result in the conversation.
