Does not match the claimed model claude-opus-5-5
行为特征与官方 claude-opus-5-5 不符,最接近官方 qwen3.7-max。
- Providers
- devcorp.me
- Claimed model
- claude-opus-5-5
- Plan
- Full check
- Reference
- 101 official models
- Requests
- 18
- Duration
- 6 分钟
01Identity check
Behaviour compared with each official referenceAnswer by answer6 answers: 0 match claude-opus-5-5;another 4 look more like qwen3.7-max;another 2 match no model clearly on their own
- qwen3.7-max>99%
- qwen3.7-plus<1%
- claude-opus-4-5<1%
- qwen3.6-max-preview<1%
- mimo-v2.5<1%
02Test details
17 items · 11 passed · 3 notes · 3 failed| No. | Check | Result | Methodology |
|---|---|---|---|
| Connectivity & auth | |||
| A-01 | Endpoint reachable | Passed | The endpoint responds normally |
| A-02 | Key authentication | Passed | API key is valid and allowed to make calls |
| A-03 | Model on sale | Passed | The model list includes claude-opus-5-5 |
| API conformance | |||
| B-01 | Basic call | Passed | Returns the expected content as instructed |
| B-02 | Response structure | Meets the spec | All OpenAI-compatible fields are present |
| B-03 | Model ID echoed back | Match | The model ID in the response matches the request |
| B-04 | Usage record | Complete and consistent | Input, output and total are all present, and the total equals the sum |
| B-05 | Output cleanliness | Nothing added | No inserted ads, watermarks or extra notes found |
| Identity check | |||
| C-01 | Independent sampling | 6 / 6 valid | 首轮未能判定,已追加一轮;独立采样均可用于比对 |
| C-02 | Brand consistency | Mismatch | Behavior does not belong to the Anthropic Claude family; brand match score 0 |
| C-03 | Model consistency | Mismatch | Model match score 0; closest to the official qwen3.7-max |
| C-04 | Cross-brand swap check | Substitution detected | Behavior is closest to qwen3.7-max from another brand |
| Billing & capabilities | |||
| D-01 | Cache billing | Does not work | The vendor offers a cache price for this model, but none of the repeated requests were billed at it. |
| D-02 | Context length | 32K failed | 32K failed: Input truncated; about 12K actually received. |
| D-03 | Streaming output | Normal | Streamed content is complete |
| D-04 | Tool use | Supported | Returns the correct function name and arguments |
| D-05 | Structured output | Did not follow the schema | The response did not follow the required JSON schema |
The same long content was sent 5 times: the 1st is billed at the regular price, and from the 2nd on the repeated part should be billed at the cache price.
The vendor offers a cache price for this model, but none of the repeated requests were billed at it.
Long documents are sent tier by tier to check whether the whole document is read.
32K failed: Input truncated; about 12K actually received.
- 32KInput truncated; about 12K actually received
Token usage check
Reported usage is used; estimates appear only when none is reported.
Method: independent samples from the tested endpoint are compared against reference samples collected from official channels; billing, context and capability items come from real calls. API keys, request bodies and raw model output are not stored.
This report reflects the endpoint at the time of testing.
Post an anonymous comment
Before you buy AI, check Mozhenzhenclaude-opus-5-5 verification report
Does not match the claimed model claude-opus-5-5
行为特征与官方 claude-opus-5-5 不符,最接近官方 qwen3.7-max。
- Providers
- devcorp.me
- Claimed model
- claude-opus-5-5
- Plan
- Full check
- Reference
- 101 official models
- Requests
- 18
- Duration
- 6 分钟
01Identity check
Behaviour compared with each official referenceAnswer by answer6 answers: 0 match claude-opus-5-5;another 4 look more like qwen3.7-max;another 2 match no model clearly on their own
- qwen3.7-max>99%
- qwen3.7-plus<1%
- claude-opus-4-5<1%
- qwen3.6-max-preview<1%
- mimo-v2.5<1%
02Test details
17 items · 11 passed · 3 notes · 3 failed| No. | Check | Result | Methodology |
|---|---|---|---|
| Connectivity & auth | |||
| A-01 | Endpoint reachable | Passed | The endpoint responds normally |
| A-02 | Key authentication | Passed | API key is valid and allowed to make calls |
| A-03 | Model on sale | Passed | The model list includes claude-opus-5-5 |
| API conformance | |||
| B-01 | Basic call | Passed | Returns the expected content as instructed |
| B-02 | Response structure | Meets the spec | All OpenAI-compatible fields are present |
| B-03 | Model ID echoed back | Match | The model ID in the response matches the request |
| B-04 | Usage record | Complete and consistent | Input, output and total are all present, and the total equals the sum |
| B-05 | Output cleanliness | Nothing added | No inserted ads, watermarks or extra notes found |
| Identity check | |||
| C-01 | Independent sampling | 6 / 6 valid | 首轮未能判定,已追加一轮;独立采样均可用于比对 |
| C-02 | Brand consistency | Mismatch | Behavior does not belong to the Anthropic Claude family; brand match score 0 |
| C-03 | Model consistency | Mismatch | Model match score 0; closest to the official qwen3.7-max |
| C-04 | Cross-brand swap check | Substitution detected | Behavior is closest to qwen3.7-max from another brand |
| Billing & capabilities | |||
| D-01 | Cache billing | Does not work | The vendor offers a cache price for this model, but none of the repeated requests were billed at it. |
| D-02 | Context length | 32K failed | 32K failed: Input truncated; about 12K actually received. |
| D-03 | Streaming output | Normal | Streamed content is complete |
| D-04 | Tool use | Supported | Returns the correct function name and arguments |
| D-05 | Structured output | Did not follow the schema | The response did not follow the required JSON schema |
The same long content was sent 5 times: the 1st is billed at the regular price, and from the 2nd on the repeated part should be billed at the cache price.
The vendor offers a cache price for this model, but none of the repeated requests were billed at it.
Long documents are sent tier by tier to check whether the whole document is read.
32K failed: Input truncated; about 12K actually received.
- 32KInput truncated; about 12K actually received
Token usage check
Reported usage is used; estimates appear only when none is reported.
Method: independent samples from the tested endpoint are compared against reference samples collected from official channels; billing, context and capability items come from real calls. API keys, request bodies and raw model output are not stored.
This report reflects the endpoint at the time of testing.