claude-opus-4-8 standard verification model-consistency report
Report ID MZZR20260826143820095DC3Generated at 2026/08/26 22:38
Matched with issues
claude-opus-4-8 standard verification result: Matched with issues (80/100). Main deductions: Model-code anchoring was unstable; Model-family identification was unstable.
80Consistency score
主要扣分项:模型代码锚定不稳定;模型族识别不稳定。
EvaluatorMozhenzhen
Tested modelClaude Opus 4.8
ProvidersIKunCode
Samples20 probes
Duration1 min 24 sec
Verification versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage34.0KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate855Precisely estimated from this test suite when no official baseline is available.Variance assessment3882% aboveMeasured token usage is above the estimate; review billing or provider logs.Usage conclusionAbove estimateThis reflects usage observability, not the final billed amount.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
claude-opus-4-8 standard verification result: Matched with issues (80/100). Main deductions: Model-code anchoring was unstable; Model-family identification was unstable.
EvaluatorMozhenzhen
Tested modelClaude Opus 4.8
ProvidersIKunCode
Samples20
Duration1 min
Test versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage34.0KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate855Precisely estimated from this test suite when no official baseline is available.Variance assessment3882% aboveMeasured token usage is above the estimate; review billing or provider logs.Usage conclusionAbove estimateThis reflects usage observability, not the final billed amount.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.