minimax-m3 deep verification model-consistency report
Report ID MZZR20260812024617033867Generated at 2026/08/12 10:46
High risk
minimax-m3 deep verification result: High risk (55/100). Main deductions: Logic-grid reasoning was unstable; Upstream response timed out; Hidden-prompt boundary was unclear.
55Consistency score
主要扣分项:逻辑网格推理不稳定;上游响应超时;隐藏提示词边界不清。
EvaluatorMozhenzhen
Tested modelMiniMax-M3
Providersnextapi.store
Samples24 probes
Duration3 min 21 sec
Verification versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage23.6KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate27.4KPrecisely estimated from this test suite when no official baseline is available.Variance assessment14% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
Verified up to 32K
All tested context tiers (32K) returned reliably.
Reference windowReference window 32KModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
minimax-m3 deep verification result: High risk (55/100). Main deductions: Logic-grid reasoning was unstable; Upstream response timed out; Hidden-prompt boundary was unclear.
EvaluatorMozhenzhen
Tested modelMiniMax-M3
Providersnextapi.store
Samples24
Duration3 min
Test versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage23.6KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate27.4KPrecisely estimated from this test suite when no official baseline is available.Variance assessment14% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
Verified up to 32K
All tested context tiers (32K) returned reliably.
Reference windowReference window 32KModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.