step-3.5-flash deep verification model-consistency report
Report ID MZZR2026070905575640DA90Generated at 2026/07/09 13:57
Doubtful
step-3.5-flash deep verification result: Doubtful (70/100). Main deductions: Logic-grid reasoning was unstable; Quota or rate limit reached; Several key evidence items need review.
70Consistency score
主要扣分项:逻辑网格推理不稳定;额度或频率限制;多个关键证据项需要复核。
EvaluatorMozhenzhen
Tested modelStep 3.5 Flash
ProvidersInfistar
Samples107 probes
Duration29 sec
Verification versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage1.6KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate82.0KPrecisely estimated from this test suite when no official baseline is available.Variance assessment98% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 5 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K 需要关注,原因:额度或频率限制。
Reference windowReference window 32KModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
A newer report is availableThis report has been superseded by a later retest. View the latest report。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
step-3.5-flash deep verification result: Doubtful (70/100). Main deductions: Logic-grid reasoning was unstable; Quota or rate limit reached; Several key evidence items need review.
EvaluatorMozhenzhen
Tested modelStep 3.5 Flash
ProvidersInfistar
Samples107
Duration29 sec
Test versionV11
Identity verdictNot verified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage1.6KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate82.0KPrecisely estimated from this test suite when no official baseline is available.Variance assessment98% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 5 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K 需要关注,原因:额度或频率限制。
Reference windowReference window 32KModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
A newer report is availableThis report has been superseded by a later retest. View the latest report。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.