Evidence before scale
Neul Labs uses a simple rule for engineering and public claims: define the decision, bound the system, preserve the baseline, measure the change and state what the evidence cannot establish. The method applies to client work, public repositories, benchmarks, R&D proposals and partnership claims.
Six stages from question to claim
The useful unit is a traceable decision record, not a dashboard screenshot or isolated headline number.
- Stage 1
Question
State the decision the evidence should change, the target user or system and what a useful negative result would mean.
- Stage 2
Boundary
Freeze versions, inputs, environment, permissions, data provenance, cost ceiling and the behaviours inside and outside scope.
- Stage 3
Baseline
Record the current system before intervention, including correctness, failure cases and resource use relevant to the decision.
- Stage 4
Intervention
Change one coherent variable or work package and keep the fallback or comparison path available.
- Stage 5
Verification
Repeat the workload, inspect outliers, validate outputs and record failures, variance and operational trade-offs.
- Stage 6
Publication
Separate observation from interpretation, link the source artefact and state versions, dates, limitations and claim owner.
Every claim has a different proof source
A repository can support an implementation claim. It cannot prove a customer outcome or confer a provider credential.
| Claim class | Primary proof | Example |
|---|---|---|
| Company fact | Public register or controlled company record | Legal name, incorporation date and registered office |
| Implementation fact | Versioned source, test or release artefact | An interface, feature or supported platform in a named version |
| Measured result | Reproducible workload, raw results and comparison method | Latency, throughput, memory, task success or compute use |
| External status | Issuer decision or public provider page | Partner, marketplace, certification, grant or programme acceptance |
| Customer outcome | Customer-approved evidence and attribution | Deployment, savings, adoption or testimonial |
| Research conclusion | Defined method, observations, uncertainty and limitations | A bounded finding that does not overgeneralise beyond the experiment |
Minimum reproducibility fields
A reader should be able to decide whether the workload resembles their own and whether rerunning it is practical.
- Source commit and dependency or model versions.
- Hardware, operating system, runtime and important configuration.
- Input provenance, size, distribution and preprocessing.
- Setup, cache state, warm-up, concurrency and number of runs.
- Correctness oracle and any allowed output tolerance.
- Raw observations, summary statistic and variance or uncertainty.
- Resource measures such as CPU, memory, I/O, tokens or GPU hours.
- Known bottlenecks, excluded cases and conflicts of interest.
AI system evidence
Combine deterministic checks, representative tasks, tool-call validation, human review and bounded model-based scoring. Record prompt, model, tool and policy versions and test adverse paths.
Agent evaluation research →R&D evidence
Separate technical uncertainty from ordinary delivery, define the knowledge gap and additionality, record negative results and reconcile each cost to a single funding source.
R&D collaboration shape →External recognition
Applications, accounts and compatibility are recorded internally. Public partner, directory, certification, grant and badge claims wait for issuer confirmation and permitted wording.
Recognition policy →Questions people ask
What does evidence before scale mean?
It means agreeing the decision, baseline and acceptance method before expanding implementation or making a broad claim. A small reversible test can reveal that the proposed architecture, provider or optimisation is not the right one.
How does Neul Labs report benchmark results?
A defensible report identifies source and dependency versions, hardware and software environment, input data, setup and warm-up, number of runs, summary statistic, correctness checks, raw results where publishable and known limits. Results are not presented as universal guarantees.
How are partner and certification claims verified?
An account, application, course, compatibility test or private conversation is not treated as awarded status. The exact public claim is made only when the issuing organisation confirms it and its branding or directory rules allow publication.
Can unsuccessful experiments be useful?
Yes. A valid negative or inconclusive result can prevent an expensive rewrite, identify a missing data or evaluation requirement, narrow the product or produce a better research question. The method should define that outcome in advance.