gpt-5.5 deep verification model-consistency report
Report ID MZZR20260805074010EE6149Generated at 2026/08/05 15:40
Matched with issues
gpt-5.5 deep verification result: Matched with issues (78/100). Main deductions: API authentication or permission error; Failed evidence items were found.
78Consistency score
主要扣分项:接口鉴权或权限异常;存在Failed证据项。
EvaluatorMozhenzhen
Tested modelGPT-5.5
Providers聚光AI
Samples119 probes
Duration49 sec
Verification versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage13.6KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate74.9KPrecisely estimated from this test suite when no official baseline is available.Variance assessment82% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 1 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K Failed,原因:接口鉴权或权限异常。
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
Scoring breakdown
Matched with issues78/100
主要扣分项:接口鉴权或权限异常;存在Failed证据项。
Callability and model visibility2 / 10
接口鉴权或权限异常(标准题:隐藏提示词边界) -8
Protocol and response structure8 / 12
接口鉴权或权限异常(标准题:隐藏提示词边界) -4
Identity and model match18 / 18
No deductions
Core capabilities and model-specific tests25 / 25
No deductions
Usage and billing consistency10 / 10
No deductions
Reliability and availability15 / 25
接口鉴权或权限异常(标准题:隐藏提示词边界) -8存在Failed证据项 -2
Test details
Expand each check to view the redacted evidence notes.
22 checks
Model list visibility
gpt-5.5 is visible in the API model list.
Passed
Test evidence
Check: Model list visibility
Status: Passed
Samples: 109
API path: /v1/models
Visible models: 109
Basic call probe
The basic API call succeeded and returned a recognizable response.
Check: Standard test: usage observability
Status: Skipped
Samples: 0
Test type: Text test
Needs review原因:接口鉴权或权限异常
跳过原因:service_failure_circuit_breaker
Check: Model-family test: series and tier identification
Status: Skipped
Samples: 0
Test type: Text test
Needs review原因:接口鉴权或权限异常
跳过原因:service_failure_circuit_breaker
Check: Deep model test: family-specific boundary
Status: Skipped
Samples: 0
Test type: Text test
Needs review原因:接口鉴权或权限异常
跳过原因:service_failure_circuit_breaker
Specialty test: token cache usage over 5 rounds
执行 专项题,用于形成模真真验真证据。 检测过程中确认服务类故障,剩余样本和后续检测项已熔断跳过。
Failed
Test evidence
Check: Specialty test: token cache usage over 5 rounds
Status: Failed
Samples: 1
Test type: Text test
Specialty: Token cache usage test (may affect real billing)
Needs review原因:接口鉴权或权限异常
Results: Passed 0 / Needs review 0 / Failed 1
Specialty test: 32K context capability
32K 长度样本未形成稳定结果,本次按边界样本记录。
Failed
Test evidence
Check: Specialty test: 32K context capability
Status: Failed
Samples: 1
Test type: Text test
Specialty: Context test
Context tier: 32K
Needs review原因:接口鉴权或权限异常
Results: Passed 0 / Needs review 0 / Failed 1
gpt-5.5 deep verification result: Matched with issues (78/100). Main deductions: API authentication or permission error; Failed evidence items were found.
EvaluatorMozhenzhen
Tested modelGPT-5.5
Providers聚光AI
Samples119
Duration49 sec
Test versionV11
Identity verdictVerified
Verification basisMozhenzhen versioned baseline
Token usage
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Measured token usage13.6KFrom the response usage fields; this is the report's primary usage measure.Test-suite estimate74.9KPrecisely estimated from this test suite when no official baseline is available.Variance assessment82% belowMeasured token usage is below the estimate, possibly due to tokenizer, compression, or usage-reporting differences.Usage conclusionBelow estimateThis reflects usage observability, not the final billed amount.
Token cache usage test
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
Tested 1 times
Observed cache share--
暂未观测到用量记录
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
Observed tokens0Cached tokens0Non-cached tokens0Average per run0
Cache share by testRepresents the API-reported cache share, not the final billing discount.
Long-context test
Shows stable responses at different context lengths and the model center reference window.
1 samples
边界在 32K
32K Failed,原因:接口鉴权或权限异常。
Reference windowReference window 1MModel-center context window; available lengths may vary by provider.Tested context lengths32KTests begin at the longest generated context and step down to verify stable, correct responses.
Current guidance
32K 档暂不作为当前线路的稳定上下文参考。
Model authenticity score
Review API availability, model identity, response completeness, and other checks separately.
Scoring breakdown
Matched with issues78/100
主要扣分项:接口鉴权或权限异常;存在Failed证据项。
Callability and model visibility2 / 10
接口鉴权或权限异常(标准题:隐藏提示词边界) -8
Protocol and response structure8 / 12
接口鉴权或权限异常(标准题:隐藏提示词边界) -4
Identity and model match18 / 18
No deductions
Core capabilities and model-specific tests25 / 25
No deductions
Usage and billing consistency10 / 10
No deductions
Reliability and availability15 / 25
接口鉴权或权限异常(标准题:隐藏提示词边界) -8存在Failed证据项 -2
Test details
Expand each check to view the redacted evidence notes.
22 checks
Model list visibility
gpt-5.5 is visible in the API model list.
Passed
Test evidence
Check: Model list visibility
Status: Passed
Samples: 109
API path: /v1/models
Visible models: 109
Basic call probe
The basic API call succeeded and returned a recognizable response.
Check: Standard test: usage observability
Status: Skipped
Samples: 0
Test type: Text test
Needs review原因:接口鉴权或权限异常
跳过原因:service_failure_circuit_breaker
Check: Model-family test: series and tier identification
Status: Skipped
Samples: 0
Test type: Text test
Needs review原因:接口鉴权或权限异常
跳过原因:service_failure_circuit_breaker
Check: Deep model test: family-specific boundary
Status: Skipped
Samples: 0
Test type: Text test
Needs review原因:接口鉴权或权限异常
跳过原因:service_failure_circuit_breaker
Specialty test: token cache usage over 5 rounds
执行 专项题,用于形成模真真验真证据。 检测过程中确认服务类故障,剩余样本和后续检测项已熔断跳过。
Failed
Test evidence
Check: Specialty test: token cache usage over 5 rounds
Status: Failed
Samples: 1
Test type: Text test
Specialty: Token cache usage test (may affect real billing)
Needs review原因:接口鉴权或权限异常
Results: Passed 0 / Needs review 0 / Failed 1
Specialty test: 32K context capability
32K 长度样本未形成稳定结果,本次按边界样本记录。
Failed
Test evidence
Check: Specialty test: 32K context capability
Status: Failed
Samples: 1
Test type: Text test
Specialty: Context test
Context tier: 32K
Needs review原因:接口鉴权或权限异常
Results: Passed 0 / Needs review 0 / Failed 1