Skip to content

Watch completebenchmark runs.

Each of these is a real session captured end to end: same prompts, same models, same passing criteria, against a GPU or hosted-API baseline. Where a clip is sped up, the badge says so.

Watch

Repeatability, concurrency, and cost, in that order.

Start with repeatability. It is the one that surprises people, because most inference stacks quietly give you a different answer under load.

Deterministic, repeatable output in live Cursor sessions

The same coding session run repeatedly produces identical output, at full throughput. GPU output varies under higher load and batching.

Hardware
Score on Xeon versus a GPU baseline
Workload
Live Cursor coding sessions, repeated

8 concurrent coding sessions versus 2 on GPU

Score-SDK holds eight concurrent Cursor-like coding sessions on a single CPU server where the GPU baseline manages two.

Hardware
One CPU server versus a GPU baseline
Workload
Eight concurrent coding sessions

3× more concurrent knowledge-search sessions

Six concurrent RAG knowledge-search sessions on Score versus two on the GPU baseline.

Hardware
Score on Xeon versus a GPU baseline
Workload
Six concurrent knowledge-search sessions

Fast, cost-efficient RAG agents on Score Engine

Concurrent retrieval-augmented agents running against a self-hosted open-weight model.

Hardware
Score Engine, self-hosted
Workload
Concurrent retrieval-augmented agents

RAG cost: self-hosted gpt-oss-120B on Gen4 Xeon vs Amazon Bedrock

Direct cost comparison for the same retrieval workload, self-hosted against a hosted API.

Model
gpt-oss-120B
Hardware
Gen4 Xeon versus Amazon Bedrock
Workload
Retrieval-augmented generation, cost comparison

Coding agent: Gemma 4-31B on Gen6 Xeon vs OpenRouter

Multi-turn coding agent, self-hosted with Score-SDK, against the same model served by a hosted API.

Model
Gemma 4-31B
Hardware
Gen6 Xeon versus OpenRouter
Workload
Multi-turn coding agent, cost comparison

Want this run on your workload instead?

We'll record the same comparison using your models, your context lengths and your concurrency, and hand you the numbers either way.