Skip to content
Engineering method · Updated 17 August 2026

Evidence before scale

Neul Labs uses a simple rule for engineering and public claims: define the decision, bound the system, preserve the baseline, measure the change and state what the evidence cannot establish. The method applies to client work, public repositories, benchmarks, R&D proposals and partnership claims.

Six stages from question to claim

The useful unit is a traceable decision record, not a dashboard screenshot or isolated headline number.

  1. Stage 1

    Question

    State the decision the evidence should change, the target user or system and what a useful negative result would mean.

  2. Stage 2

    Boundary

    Freeze versions, inputs, environment, permissions, data provenance, cost ceiling and the behaviours inside and outside scope.

  3. Stage 3

    Baseline

    Record the current system before intervention, including correctness, failure cases and resource use relevant to the decision.

  4. Stage 4

    Intervention

    Change one coherent variable or work package and keep the fallback or comparison path available.

  5. Stage 5

    Verification

    Repeat the workload, inspect outliers, validate outputs and record failures, variance and operational trade-offs.

  6. Stage 6

    Publication

    Separate observation from interpretation, link the source artefact and state versions, dates, limitations and claim owner.

Every claim has a different proof source

A repository can support an implementation claim. It cannot prove a customer outcome or confer a provider credential.

Claim classPrimary proofExample
Company factPublic register or controlled company recordLegal name, incorporation date and registered office
Implementation factVersioned source, test or release artefactAn interface, feature or supported platform in a named version
Measured resultReproducible workload, raw results and comparison methodLatency, throughput, memory, task success or compute use
External statusIssuer decision or public provider pagePartner, marketplace, certification, grant or programme acceptance
Customer outcomeCustomer-approved evidence and attributionDeployment, savings, adoption or testimonial
Research conclusionDefined method, observations, uncertainty and limitationsA bounded finding that does not overgeneralise beyond the experiment
Benchmark record

Minimum reproducibility fields

A reader should be able to decide whether the workload resembles their own and whether rerunning it is practical.

  • Source commit and dependency or model versions.
  • Hardware, operating system, runtime and important configuration.
  • Input provenance, size, distribution and preprocessing.
  • Setup, cache state, warm-up, concurrency and number of runs.
  • Correctness oracle and any allowed output tolerance.
  • Raw observations, summary statistic and variance or uncertainty.
  • Resource measures such as CPU, memory, I/O, tokens or GPU hours.
  • Known bottlenecks, excluded cases and conflicts of interest.

AI system evidence

Combine deterministic checks, representative tasks, tool-call validation, human review and bounded model-based scoring. Record prompt, model, tool and policy versions and test adverse paths.

Agent evaluation research →

R&D evidence

Separate technical uncertainty from ordinary delivery, define the knowledge gap and additionality, record negative results and reconcile each cost to a single funding source.

R&D collaboration shape →

External recognition

Applications, accounts and compatibility are recorded internally. Public partner, directory, certification, grant and badge claims wait for issuer confirmation and permitted wording.

Recognition policy →

Questions people ask

What does evidence before scale mean?

It means agreeing the decision, baseline and acceptance method before expanding implementation or making a broad claim. A small reversible test can reveal that the proposed architecture, provider or optimisation is not the right one.

How does Neul Labs report benchmark results?

A defensible report identifies source and dependency versions, hardware and software environment, input data, setup and warm-up, number of runs, summary statistic, correctness checks, raw results where publishable and known limits. Results are not presented as universal guarantees.

How are partner and certification claims verified?

An account, application, course, compatibility test or private conversation is not treated as awarded status. The exact public claim is made only when the issuing organisation confirms it and its branding or directory rules allow publication.

Can unsuccessful experiments be useful?

Yes. A valid negative or inconclusive result can prevent an expensive rewrite, identify a missing data or evaluation requirement, narrow the product or produce a better research question. The method should define that outcome in advance.

What decision should the evidence change?

Send the current system, the proposed change, the relevant workload and the threshold that would alter your decision.

admin@neullabs.com