deepseek-v4-flash deep verification model-consistency report
Report ID MZZR20260708074256EF43F3Generated at 2026/07/08 15:42
Matched with issues
deepseek-v4-flash deep verification result: Matched with issues (84/100). Main deductions: API service error; Several key evidence items need review.
84Consistency score
主要扣分项:接口服务异常;多个关键证据项需要复核。
EvaluatorMozhenzhen
Tested modelDeepSeek V4 Flash
ProvidersInfistar
Samples107 probes
Duration8 min 33 sec
Verification versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage918From the response usage fields; this is the report's primary usage measure.Test-suite estimate1.7MPrecisely estimated from this test suite when no official baseline is available.Variance assessment100% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 5 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K 需要关注,原因:接口服务异常。
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
A newer report is availableThis report has been superseded by a later retest. View the latest report。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
Scoring breakdown
Matched with issues84/100
主要扣分项:接口服务异常;多个关键证据项需要复核。
Callability and model visibility10 / 10
No deductions
Protocol and response structure8 / 12
接口服务异常(深度题:逻辑网格推理、深度题:复测一致性) -4
Identity and model match18 / 18
No deductions
Core capabilities and model-specific tests25 / 25
No deductions
Usage and billing consistency10 / 10
No deductions
Reliability and availability13 / 25
接口服务异常(深度题:逻辑网格推理、深度题:复测一致性) -8多个关键证据项需要复核 -4
Test details
Expand each check to view the redacted evidence notes.
18 checks
Model list visibility
deepseek-v4-flash is visible in the API model list.
Passed
Test evidence
Check: Model list visibility
Status: Passed
Samples: 85
API path: /v1/models
Visible models: 85
Basic call probe
The basic API call succeeded and returned a recognizable response.
deepseek-v4-flash deep verification result: Matched with issues (84/100). Main deductions: API service error; Several key evidence items need review.
EvaluatorMozhenzhen
Tested modelDeepSeek V4 Flash
ProvidersInfistar
Samples107
Duration9 min
Test versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage918From the response usage fields; this is the report's primary usage measure.Test-suite estimate1.7MPrecisely estimated from this test suite when no official baseline is available.Variance assessment100% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 5 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K 需要关注,原因:接口服务异常。
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
A newer report is availableThis report has been superseded by a later retest. View the latest report。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
Scoring breakdown
Matched with issues84/100
主要扣分项:接口服务异常;多个关键证据项需要复核。
Callability and model visibility10 / 10
No deductions
Protocol and response structure8 / 12
接口服务异常(深度题:逻辑网格推理、深度题:复测一致性) -4
Identity and model match18 / 18
No deductions
Core capabilities and model-specific tests25 / 25
No deductions
Usage and billing consistency10 / 10
No deductions
Reliability and availability13 / 25
接口服务异常(深度题:逻辑网格推理、深度题:复测一致性) -8多个关键证据项需要复核 -4
Test details
Expand each check to view the redacted evidence notes.
18 checks
Model list visibility
deepseek-v4-flash is visible in the API model list.
Passed
Test evidence
Check: Model list visibility
Status: Passed
Samples: 85
API path: /v1/models
Visible models: 85
Basic call probe
The basic API call succeeded and returned a recognizable response.