mimo-v2.5 deep verification model-consistency report
Report ID MZZR20260826033428E67D04Generated at 2026/08/26 11:34
Doubtful
mimo-v2.5 deep verification result: Doubtful (69/100). Main deductions: Logic-grid reasoning was unstable; Hidden-prompt boundary was unclear; Usage fields were not observable.
69Consistency score
主要扣分项:逻辑网格推理不稳定;隐藏提示词边界不清;usage 用量字段不可观测。
EvaluatorMozhenzhen
Tested modelMiMo-V2.5
Providersinfistar.cc
Samples130 probes
Duration3 min 39 sec
Verification versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage675.9KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate1.7MPrecisely estimated from this test suite when no official baseline is available.Variance assessment60% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 5 times
Observed cache share93.5%
缓存字段表现稳定
多次测试均返回缓存 Token,说明这条链路的长对话用量记录有参考价值。
Observed tokens15.4KCached tokens14.4KNon-cached tokens1.0KAverage per run3.1K
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
Verified up to 1M
All tested context tiers (1M) returned reliably.
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths1MTests begin at the longest generated context and step down to verify stable, correct responses.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
mimo-v2.5 deep verification result: Doubtful (69/100). Main deductions: Logic-grid reasoning was unstable; Hidden-prompt boundary was unclear; Usage fields were not observable.
EvaluatorMozhenzhen
Tested modelMiMo-V2.5
Providersinfistar.cc
Samples130
Duration4 min
Test versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage675.9KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate1.7MPrecisely estimated from this test suite when no official baseline is available.Variance assessment60% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 5 times
Observed cache share93.5%
缓存字段表现稳定
多次测试均返回缓存 Token,说明这条链路的长对话用量记录有参考价值。
Observed tokens15.4KCached tokens14.4KNon-cached tokens1.0KAverage per run3.1K
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
Verified up to 1M
All tested context tiers (1M) returned reliably.
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths1MTests begin at the longest generated context and step down to verify stable, correct responses.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.