Technical foundation working today
Exercised in the synthetic demonstrator and evidenced by its test suite. Says nothing about behaviour on real data.
- Synthetic event ingestion
- Seeded scenarios producing reproducible event streams with a published fixture hash.
- Deterministic state projection
- State rebuilt from events on every read; any point in a run can be reconstructed exactly.
- Current detector rules
- Deterministic detectors for delay, missing prerequisite, stale dependency, non-response and conflict.
- Authority policy enforcement
- Deny by default; clinical and uncatalogued requests refused and recorded.
- Human approval mechanics
- Single and dual approval, separation of duties, amendment, rejection and expiry.
- Permitted action execution
- Idempotent execution through simulated connectors in the synthetic environment.
- Outcome monitoring registration
- Expected outcomes tracked, bounded retry, compensation and escalation.
- Tamper evident audit verification
- Hash-chained records with a verification routine and a tamper demonstration.
Questions to be proven
Several have real negative answers available. Finding one would be a reason to stop, and we would rather find it early.
Bounded AI reasoning on heterogeneous NHS data
high uncertaintyNothing in the current build calls a model. The architecture reserves a bounded, validated place for one, and that place is empty.
Evaluation. Head to head against the deterministic baseline on a locked corpus, measuring incremental value net of added error, cost and latency.
Precision and false positive rates in real workflows
high uncertaintyA detector that is precise on clean synthetic events may be worse than useless on data that arrives late, out of order and contradictory.
Evaluation. Precision, recall and F1 per constraint type against a labelled corpus, with thresholds pre-registered before the run.
Generalisation across pathways
moderate uncertaintyIf most of what differs between pathways and organisations needs code rather than configuration, this becomes consultancy and does not scale.
Evaluation. The same core deployed against multiple configurations, measuring the proportion of variation requiring code change.
NHS integration and identity assurance
high uncertaintyWhich systems hold the relevant state, what they expose, and who is authenticated when an approval is given.
Evaluation. Integration conformance testing and state reconciliation accuracy against a source of truth.
Operational adoption and human override behaviour
moderate uncertaintyAn approval control that is rubber-stamped under workload provides the appearance of a boundary without the substance. No real user has used it.
Evaluation. Observed task studies measuring approval latency and error rate under realistic load, including proposals that should be rejected.
Real productivity effect and attribution
high uncertaintyWhether closing the execution gap changes anything measurable in a live service, against a moving baseline and concurrent initiatives.
Evaluation. Prospective evaluation with a declared comparator, separating cashable saving from released capacity from theoretical value.
Safety, security and regulatory assurance
mixed uncertaintyClinical safety classification and information governance carry genuine uncertainty because obligations follow from an intended-use determination not yet made.
Evaluation. Specialist assessment against each applicable standard, with independent testing where required.
Evaluation approach
Comparators, success measures and failure criteria are specified before a run, not chosen afterwards from whatever the data supported. Safety control failures are release blockers rather than averages. Failures and subgroup results are published alongside successes.
Where genuine uncertainty does not exist, we say so. The deterministic execution loop and the audit chain are engineering rather than research, and dressing them up as research questions would be the fastest way to lose a technical assessment.
Three perspectives we need
We are early enough that outside judgement still changes what gets built. That is the point of asking now rather than after two years of engineering.
NHS operational teams
Chief operating officers, elective recovery leads, transformation directors, and operational productivity teams.
Why this perspective
- The execution gap we are building around is drawn from reasoning, not from completed discovery. It may be wrong, or real but immaterial.
- The workflow in the demonstrator is our invention. Whether it resembles how your service actually operates is something only you can tell us.
- The narrowest, most repeated, most worth-doing task is the one that matters, and it is almost certainly not the one an outsider would guess.
What we would ask you
- 01Whether the operational problem shown is material and frequent in services you know.
- 02Which part of the demonstrated workflow is least realistic or incomplete.
- 03Which administrative actions could safely be prepared automatically, and which must always wait for a person.
- 04What would need resolving before your organisation would consider any evaluation.
Clinical safety and governance
Clinical Safety Officers, information governance and data protection professionals, security specialists and responsible AI leads.
Why this perspective
- The boundary between administrative coordination and clinical judgement is drawn by us, and it is the most consequential design decision in the system.
- Intended use determines regulatory obligation. We have deliberately not asserted that Leva Health sits outside medical device regulation, because that is a specialist determination and not ours to make.
- An approval control that gets rubber-stamped under real workload provides the appearance of a safety boundary without the substance, and no real user has yet used ours.
What we would ask you
- 01Whether the five authority classes are drawn in the right places, and what is misclassified.
- 02Which actions should never be automated regardless of approval.
- 03What the intended-use statement is missing, and what it wrongly assumes.
- 04What evidence your organisation would require before any pilot with real data.
Research and evaluation
Health economists, implementation researchers, statisticians and independent evaluators.
Why this perspective
- Attributing operational change in a live service, against a moving baseline and concurrent improvement initiatives, is a genuine methodological problem and not one we should solve alone.
- Our economic model rests on stated assumptions with no evidential basis. They are labelled as such, and they need replacing with something defensible.
- Cashable saving, released capacity and theoretical value are three different things, and conflating them is the most common way health technology overstates its case.
What we would ask you
- 01What a credible comparator looks like for this kind of operational intervention.
- 02Which outcome measures would be meaningful rather than merely available.
- 03What failure criteria should be declared in advance, so a null result is recognisable as one.
- 04Where our assumption register is weakest.
Take a look
Review of the synthetic demonstrator is by individual invitation and takes about 25 minutes: a short orientation, a guided run, then a structured questionnaire. Responses are used as discovery evidence and are never presented as validation of the product.