mimo-v2.5-pro deep verification model-consistency report
Report ID MZZR20260710085146E52159Generated at 2026/07/10 16:51
High risk
mimo-v2.5-pro deep verification result: High risk (49/100). Main deductions: Logic-grid reasoning was unstable; Hidden-prompt boundary was unclear; Tag-structure repetition was unstable.
49Consistency score
主要扣分项:逻辑网格推理不稳定;隐藏提示词边界不清;标签结构复述不稳定。
EvaluatorMozhenzhen
Tested modelMiMo-V2.5-Pro
ProvidersInfistar
Samples101 probes
Duration2 min 40 sec
Verification versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage8.8KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate6.7KPrecisely estimated from this test suite when no official baseline is available.Variance assessment32% aboveMeasured token usage is above the estimate; review billing or provider logs.Usage conclusionAbove estimateThis reflects usage observability, not the final billed amount.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
mimo-v2.5-pro deep verification result: High risk (49/100). Main deductions: Logic-grid reasoning was unstable; Hidden-prompt boundary was unclear; Tag-structure repetition was unstable.
EvaluatorMozhenzhen
Tested modelMiMo-V2.5-Pro
ProvidersInfistar
Samples101
Duration3 min
Test versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage8.8KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate6.7KPrecisely estimated from this test suite when no official baseline is available.Variance assessment32% aboveMeasured token usage is above the estimate; review billing or provider logs.Usage conclusionAbove estimateThis reflects usage observability, not the final billed amount.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.