claude-sonnet-5 deep verification model-consistency report
Report ID MZZR2026080406435294AC88Generated at 2026/08/04 14:43
Matched with issues
claude-sonnet-5 deep verification result: Matched with issues (83/100). Main deductions: Origin connection interrupted; Self-reported model identity needs review; Failed evidence items were found.
83Consistency score
主要扣分项:源站连接中断;模型身份自报待复核;存在Failed证据项。
EvaluatorMozhenzhen
Tested modelClaude Sonnet 5
Providerspomoai
Samples22 probes
Duration4 min 21 sec
Verification versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage1.8KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate74.9KPrecisely estimated from this test suite when no official baseline is available.Variance assessment98% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 1 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K Failed,原因:32K 长度样本未形成稳定结果,本次按边界样本记录。
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
Scoring breakdown
Matched with issues83/100
主要扣分项:源站连接中断;模型身份自报待复核;存在Failed证据项。
Callability and model visibility10 / 10
No deductions
Protocol and response structure8 / 12
源站连接中断(标准题:标签结构复述) -4
Identity and model match15 / 18
模型身份自报待复核(快速题:模型身份自报) -3
Core capabilities and model-specific tests25 / 25
No deductions
Usage and billing consistency10 / 10
No deductions
Reliability and availability15 / 25
源站连接中断(标准题:标签结构复述) -8存在Failed证据项 -2
Test details
Expand each check to view the redacted evidence notes.
23 checks
Model list visibility
claude-sonnet-5 is visible in the API model list.
Passed
Test evidence
Check: Model list visibility
Status: Passed
Samples: 10
API path: /v1/models
Visible models: 10
Basic call probe
The basic API call succeeded and returned a recognizable response.
claude-sonnet-5 deep verification result: Matched with issues (83/100). Main deductions: Origin connection interrupted; Self-reported model identity needs review; Failed evidence items were found.
EvaluatorMozhenzhen
Tested modelClaude Sonnet 5
Providerspomoai
Samples22
Duration4 min
Test versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage1.8KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate74.9KPrecisely estimated from this test suite when no official baseline is available.Variance assessment98% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 1 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K Failed,原因:32K 长度样本未形成稳定结果,本次按边界样本记录。
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
Scoring breakdown
Matched with issues83/100
主要扣分项:源站连接中断;模型身份自报待复核;存在Failed证据项。
Callability and model visibility10 / 10
No deductions
Protocol and response structure8 / 12
源站连接中断(标准题:标签结构复述) -4
Identity and model match15 / 18
模型身份自报待复核(快速题:模型身份自报) -3
Core capabilities and model-specific tests25 / 25
No deductions
Usage and billing consistency10 / 10
No deductions
Reliability and availability15 / 25
源站连接中断(标准题:标签结构复述) -8存在Failed证据项 -2
Test details
Expand each check to view the redacted evidence notes.
23 checks
Model list visibility
claude-sonnet-5 is visible in the API model list.
Passed
Test evidence
Check: Model list visibility
Status: Passed
Samples: 10
API path: /v1/models
Visible models: 10
Basic call probe
The basic API call succeeded and returned a recognizable response.