Rankingslive route intelligenceEmbeddings market

Embedding Model Rankings

Compare market share, route quality, task spend, benchmark fit, service tier, endpoint health, and GPU readiness in one practical ranking surface for customers, suppliers, and internal sales teams.

Market lens

Usage, spend, share

Routing signal

Quality, latency, cost

Buyer view

GPU-ready capacity

Weekly usage

38.4M

modeled requests ranked by route family

AI Token spend

$2.8M

credit share across models, apps, and routes

Ranking tracks

12

usage, spend, share, benchmarks, tools, images, apps, GPU

GPU readiness

18 lanes

serverless, reserved, dedicated, or private candidates

Window

7D live preview

Data scope

Embeddings route simulation

Export

model shortlist ready

Embedding Model Rankings

Vector search, memory, retrieval, classification, and app knowledge-base usage ranked by volume and cost.

weekly volume

166B

measurement

weekly embedded tokens

active routes

48+

Top Models

Weekly usage of embeddings routes across Aurona.

180B135B90B45B0
2025-06-30Aug 4Sep 8Oct 13Nov 17Dec 22Jan 26Mar 2Apr 6May 11Jun 15

subpage tracks

Each track changes the data table, cards, route config, and benchmark controls below.

Explore

Embeddings subpage

Market Share

Model and provider share for embeddings traffic, combining routed usage, AI Token credit spend, route retention, and fallback depth.

leader

text-embedding-3

share

30.8%

volume

126B

rankmodelshare7dvolumeroutegpu lane

Embeddings benchmarks

Benchmark scorecard

The scorecard converts raw model tests into route-ready ranking signals for product teams.

Quality89

human preference, instruction following, and route judge agreement

Retrieval fit84

RAG relevance, tool-call repair, and long-context recall

Latency83

p50 response time, queue depth, and stream start

Cost control84

AI Token margin, fallback waste, and cache efficiency

Reliability83

availability, retry success, and provider health

Embeddings performance

Latency, uptime, and capacity

p50 latency

92ms

availability

99.96%

retry save

87%

GPU reserve

L40S pool

stream start78%
provider health76%
fallback depth82%
cache hit rate91%

Embeddings economics

Task Spend

Credit consumption is grouped by workload so Aurona can rank not only models, but the businesses and APIs that create durable demand.

Aurona rankings console

Ranking views for models, routes, apps, and GPU lanes

Ranking sections are mapped into the Aurona business layer: AI Token spend, route quality, provider health, policy fit, latency, and private GPU readiness.

matched models

4

usage index

267

bench runs

142

Top Models

Animated 7D ranking for context length.

View models

selected

DeepSeek

score

95

usage

96%

latency

640ms

Model Leaderboard

Click any row to update the config, benchmark panel, and route preview.

Context Length
rankmodelscoreusagespendlatencyroute

By Task Spend

Share of AI Token credit spend by model family and workload type.

Other models

Market Share

Provider share across requests, tokens, and application adoption.

Other models

Benchmarks

Composite reasoning, instruction following, coding, vision, and safety scores.

Other models

Performance

Latency, availability, routing stability, and cost-efficiency signals.

Other models

Natural Languages

Multilingual assistants, translation, finance content, and regional applications.

Other models

Programming

Code generation, repo context, tool use, debugging, and PR review routes.

Other models

Context Length

Large-window models for documents, research, memory, and agent history.

Other models

Tool Calls

Function calling, structured JSON, multi-step tools, and production agents.

Other models

Images

Image input, vision reasoning, document images, and generated media pipelines.

Other models

Apps

Models ranked by downstream apps, app growth, and repeat customer usage.

Other models

GPU Ready

Models that can support reserved capacity, private deployments, or GPU-backed lanes.

Other models

Aurona Route Rankings

Rank models together with policy, cost, and GPU supply

Fusion

Top Apps

App rankings connect model demand to the actual products that consume Aurona credits.

Browse apps

Meter-Safe Ranking Preview · simulated

Rank one native measure at a time.

Keep tokens, images, tool calls, media time, and GPU hours in separate evidence tracks. Aurona can normalize position inside a track without converting unlike units into a false universal score.

ranking preview ready · planning only

Native unit

tokens

Window

7D · through Sep 5 UTC

Proposed endpoint

POST /v1/rankings/native-meter-preview

meter_rank_00uotq20

1

Text Balanced

public-opt-in · serverless · 2026-09-05

840B

tokens

2

Text Deliberate

public-opt-in · reserved-review · 2026-09-05

620B

tokens

Simulated evidence only. This preview does not publish market claims, convert native usage into billing, debit AI Tokens, rank private activity publicly, move traffic, or allocate GPUs.

Multimodal Evaluation Lab · simulated

Turn visual capability evidence into a route preview.

Hold candidates to the same prompt family, minimum case count, objective, and GPU capacity scope before comparing quality, generation time, and modeled AI Token use.

route preview ready · planning only

Proposed endpoint

POST /v1/evaluations/multimodal-route-preview

media_eval_00cyw2vn

Selected route preview

Image Precise

reserved-review · 12 shared cases

91%pass7.8sgeneration4.2AI Tokens

Image Precise

eligible · 12 cases · reserved-review

Image Fast

eligible · 12 cases · serverless

aurona/image-studio

excluded · capability mismatch

aurona/image-new

excluded · insufficient evaluation cases

Simulated evaluation evidence only. This preview does not run media jobs, publish benchmark claims, debit AI Tokens, move production traffic, or reserve GPUs.

Ranking Feed Lab · simulated

Make every ranking query explain its evidence.

Preview model, route, and app feeds with a declared measure, resolved window, visibility scope, stable version, portable dataset manifest, citation, AI Token context, and GPU capacity path.

3 rows · planning only

Feed preview

/v1/datasets/rankings-preview?surface=all&scope=workspace-only&metric=ai-token-volume&window=7d

Version

aurona-ranking-preview-v3

Receipt

ranking_00av2q9o

Schema

aurona-ranking-feed/3

Format

JSON manifest

Data through

Sep 1 · UTC

Resolved window

7 days

Measure

modeled AI Token volume index

1

Balanced Text

models · public-opt-in · 124,000 modeled requests

84

serverless

2

Agent Route

routes · public-opt-in · 108,000 modeled requests

71

reserved-review

3

Private Agent Route

routes · workspace-only · 91,000 modeled requests

63

dedicated-review

Portable citation

Aurona.ai, "Aurona Ranking Preview" (aurona-ranking-preview-v3), modeled AI Token volume index, trailing 7 days through 2026-09-01 UTC, workspace-only scope, simulated.

Origin: simulated Aurona route evidence. This manifest identifies the dataset and methodology; it does not state reuse rights.

Simulated ranking evidence. Growth compares adjacent complete windows and signals adoption, not model quality. This preview does not publish traffic, change visibility, debit AI Tokens, or reserve GPU capacity.