gpt-5.6-sol deep verification model-consistency report
High risk
gpt-5.6-sol 本次模型一致性深度验真结果:High risk(41/100)。主要扣分项:基础调用不可用;存在Failed证据项。
主要扣分项:基础调用不可用;存在Failed证据项。
gpt-5.6-sol 本次模型一致性深度验真结果:High risk(41/100)。主要扣分项:基础调用不可用;存在Failed证据项。
主要扣分项:基础调用不可用;存在Failed证据项。
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
No per-test records are available; only the aggregate result is shown.
Shows stable responses at different context lengths and the model center reference window.
基础调用未成功(基础调用不可用),长上下文专项已跳过。
基础调用未成功(基础调用不可用),长上下文专项已跳过。
Review API availability, model identity, response completeness, and other checks separately.
主要扣分项:基础调用不可用;存在Failed证据项。
Expand each check to view the redacted evidence notes.
gpt-5.6-sol is visible in the API model list.
Check: Model list visibility Status: Passed Samples: 4 API path: /models Visible models: 4
基础调用未成功,可能是 Key 权限、额度、接口路径或模型能力类型问题。
Check: Basic call probe Status: Failed Samples: 1 Needs review原因:基础调用不可用
执行 通用题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Quick test: model self-identification Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: instruction following Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: JSON structure Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:protocol。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: tag-structure repetition Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: hidden-prompt boundary Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: brand boundary Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:usage。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: usage observability Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Deep test: logic-grid reasoning Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:stability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Deep test: retest consistency Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 品牌题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 品牌题:OpenAI 身份与安全边界 证据边界 Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 品牌题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 深度品牌题:OpenAI 身份与安全边界 Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 模型级题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Model-specific test: model-code anchoring Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 模型级题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Model-family test: series and tier identification Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 模型级题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Deep model test: family-specific boundary Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:OpenAI Responses 原生协议 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:流式 / 非流式一致性 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:工具调用结构 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:JSON Schema 结构化输出 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
Check: Specialty test: token cache usage over 5 rounds Status: Skipped Samples: 0 Test type: Text test Specialty: Token cache usage test (may affect real billing) 跳过原因:基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
Check: Specialty test: 32K context capability Status: Skipped Samples: 0 Test type: Text test Specialty: Context test 跳过原因:基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
Be the first to leave a review.
AI service research and model verification reportgpt-5.6-sol 本次模型一致性深度验真结果:High risk(41/100)。主要扣分项:基础调用不可用;存在Failed证据项。
Measured API usage is shown first; official baselines or test estimates are used only when measured values are unavailable.
Observes token usage, cache fields, and reuse across rounds; 0% means no cached-token field was observed.
当前接口没有稳定返回 Token 用量记录,建议结合账单明细或接口日志复核。
No per-test records are available; only the aggregate result is shown.
Shows stable responses at different context lengths and the model center reference window.
基础调用未成功(基础调用不可用),长上下文专项已跳过。
基础调用未成功(基础调用不可用),长上下文专项已跳过。
Review API availability, model identity, response completeness, and other checks separately.
主要扣分项:基础调用不可用;存在Failed证据项。
Expand each check to view the redacted evidence notes.
gpt-5.6-sol is visible in the API model list.
Check: Model list visibility Status: Passed Samples: 4 API path: /models Visible models: 4
基础调用未成功,可能是 Key 权限、额度、接口路径或模型能力类型问题。
Check: Basic call probe Status: Failed Samples: 1 Needs review原因:基础调用不可用
执行 通用题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Quick test: model self-identification Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: instruction following Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: JSON structure Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:protocol。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: tag-structure repetition Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: hidden-prompt boundary Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: brand boundary Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:usage。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Standard test: usage observability Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Deep test: logic-grid reasoning Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 通用题,用于形成模真真验真证据,评分维度:stability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Deep test: retest consistency Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 品牌题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 品牌题:OpenAI 身份与安全边界 证据边界 Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 品牌题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 深度品牌题:OpenAI 身份与安全边界 Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 模型级题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Model-specific test: model-code anchoring Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 模型级题,用于形成模真真验真证据,评分维度:identity。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Model-family test: series and tier identification Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 模型级题,用于形成模真真验真证据,评分维度:ability。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: Deep model test: family-specific boundary Status: Skipped Samples: 0 Test type: Text test 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:OpenAI Responses 原生协议 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:流式 / 非流式一致性 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:工具调用结构 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
按CPA 官方产品源确定的方法执行。执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
Check: 正式来源必检:专项:JSON Schema 结构化输出 Status: Skipped Samples: 0 跳过原因:基础调用探针Failed,后续题包已跳过,避免继续消耗 token。
执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
Check: Specialty test: token cache usage over 5 rounds Status: Skipped Samples: 0 Test type: Text test Specialty: Token cache usage test (may affect real billing) 跳过原因:基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
执行 专项题,用于形成模真真验真证据。 基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
Check: Specialty test: 32K context capability Status: Skipped Samples: 0 Test type: Text test Specialty: Context test 跳过原因:基础调用探针Failed,专项检测已跳过,避免继续消耗 token。
Be the first to leave a review.