Submit a request
Pick a synthetic employee, write the request the way an employee would, and watch the engine decide what it can and cannot do with it.
Decision record
No request yet. The record shows exactly what the system read, what it cited, and who has to approve before anything happens.
Answer to the employee
Recommended action
Citations
Decision trace
Approval queue
Enter a reviewer name to record a demo decision. Reviewer identities are self-reported, not authenticated. Nothing here writes to an HRIS.
All cases
| Case | Employee | Request | Status | Risk | Reviewer |
|---|
Workflow measures
Observed from the case store in this session. Hours saved is an estimate from a fixed per-case assumption, labeled as such.
Quality baseline
Separate from workflow measures. The 60-case suite runs deterministically in CI and the committed report must match the engine. Held-out cases show how far the rules generalize.
Audit events
| Time | Case | Actor | Event | Details |
|---|
Controls and where they live in the code
These tests check specific synthetic cases, not general safety or production validation. Keyword recognition has known holdout failures; names are self-reported and logs remain in memory.
| Control | Implementation | Test |
|---|---|---|
| Known injection phrasings stop before any tool access | engine.py INJECTION patterns, checked first | test_prompt_injection_is_refused_before_any_data_access |
| Known sensitive-data requests are refused | engine.py SENSITIVE terms plus other-person detection | test_other_persons_sensitive_data_is_refused |
| Employee Relations concerns are not investigated by the system | engine.py ER_TERMS route to employee_relations | test_employee_relations_language_escalates_without_fact_finding |
| Recognized legal keywords produce Legal escalation status | engine.py LEGAL_TERMS | test_legal_language_routes_to_legal |
| Minimal synthetic persona list in bootstrap | api.py bootstrap field selection | test_bootstrap_exposes_only_minimum_employee_fields |
| Only active policy versions are cited | data.py active_policies() filters superseded | test_remote_work_uses_active_policy_version_not_superseded |
| Every consequential outcome needs a named approver | engine.py approval_required on all workflow paths | test_every_consequential_outcome_requires_approval |
| Fail closed on missing records or policy gaps | engine.py missing employee and policy_gap branches | test_missing_employee_fails_closed, test_uk_employee_with_no_regional_policy_escalates_as_policy_gap |
| Audit trail for every creation and decision | store.py _log() | test_resolve_then_approve_round_trip |
| Evidence can't drift from code | evals/run.py plus committed report | test_committed_report_matches_the_current_engine |
What this demo is not
No authentication, no persistent database, no real data. The live demo runs entirely in your browser, so each visitor gets a private session and nothing typed here leaves the page. Authentication, authorization, and data retention are release gates in docs/governance-and-risk.md. The engine is deterministic on purpose so every safety decision can be reproduced; production use is withheld after the original 0/16 holdout result. A redesign must address hidden risks and be evaluated on new unseen cases as well as the regression suite, without assuming a model is the solution.