DX-Score General is the current package. It does not include industry-specific validation.
AI safety test records
Evidence for tested AI.
DX-Score records a specific system, version, scope and result under one fixed, comparable AIDX safety test.


What is DX-Score?
A scoped AI safety test certificate for an identified model, chatbot, agent or application version.
One test result. One traceable public record.
- Traceable evidence
- Method version, test-suite version, leaderboard baseline, tested date and verification ID stay connected.
- Public result, private detail
- The public record stays concise. Case outcomes, findings and remediation evidence remain in the private report.
Why DX-Score is needed
AI safety results cannot be compared when cases, methods or versions change.
Same cases. Same method. Comparable results.
Without a locked baseline
A score can look precise while saying little about how another system would perform under the same conditions.
With a locked baseline
Rank, qualification and tested scope can be traced to a named cohort, suite version and scoring method.
How the test works
DX-Score General applies one fixed, versioned 6,000-case suite to both cohort models and customer systems.
BenchDX
3,000Tests normal-use safety across harmful content, fairness, privacy, data leakage and legal-risk categories.
RobustDX
3,000Tests resistance to instruction manipulation, goal hijacking, encoded attacks and textual perturbations.
Define the tested scope
Lock the system identity, version, configuration and intended deployment boundary before testing begins.
Run the fixed suite
AIDX executes all 6,000 cases under the declared method and test-suite versions.
Review the evidence
Automated scoring and risk-based human review produce the private test report.
Compare the cohort
Compare the result with the locked open-source and commercial model cohort.
Publish the record
A qualifying result receives a public verification ID tied to its tested scope and baseline.
What Certified means
The tested result outperforms at least 50% of the locked open-source and commercial model cohort.
The cohort and baseline version remain attached to the test record so later leaderboard changes do not rewrite the historical result.

