Aurona.ai

Models

A model marketplace built for production routing.

Search public and private model supply, compare token economics, inspect route health, and choose the right provider or GPU-backed lane through Aurona.

OpenAI-compatible

Token ledger

GPU route layer

Source models

400+

Long context

1M+

Entry routes

Free

openai/gpt-5.5-pro

price, latency, policy, fallback

anthropic/claude-opus-4.8

price, latency, policy, fallback

deepseek/deepseek-v4-flash

price, latency, policy, fallback

sovereign/private

price, latency, policy, fallback

Model control plane

The model list becomes useful when every row can route.

Aurona should treat each model as a productized supply object: price, context, modality, provider health, policy support, fallback depth, and compute path all matter.

Source catalog

600+ market models

Public, private, and partner model supply is mapped into Aurona routes with provider, context, modality, price, and endpoint signals.

Route score

0-100

Combines quality, provider health, policy fit, fallback depth, token price, and capacity readiness.

Policy gates

ZDR + region

Routes can be filtered by retention, approved vendors, parameter support, and customer tier.

Compute path

Hosted or GPU

High-volume routes can move from public provider APIs into private or reserved inference lanes.

Key mode

Aurona or BYOK

Route previews can distinguish Aurona-managed access, customer provider vault keys, and hybrid fallback.

Service tier

5 classes

Shared, priority, serverless, reserved, and dedicated routes help customers match latency and capacity to workload risk.

Endpoint health

queue + cache

Dedicated and serverless routes can expose queue depth, warm-cache state, quota, saturation, and retry reason.

Policy actions

block/mask/warn

Catalog tests can preview which route, provider, prompt, output, or budget rules would fire before production traffic is sent.

Prompt history

saved runs

Approved playground tests can carry model parameters, modality, route score, credit estimate, and redaction state into launch review.

Trace fanout

broadcast

Route traces can be copied into usage logs, webhooks, observability tools, and billing exports.

Key detail

usage + logs

Every catalog test can link back to the API key, budget progress, request filters, and route history.

Admin API

catalog to control

Model rows can become route aliases, budget rules, provider-key policies, and app meters through management endpoints.

Regional fit

sovereign

Model rows can expose region, vendor, and capacity constraints before a regulated route is approved.

Catalog lifecycle

launch to retire

Model age, preview state, version pinning, replacement route, and retirement date can be reviewed before an app depends on a slug.

Output use

training review

Provider-declared training or distillation eligibility can be surfaced as review metadata instead of being inferred from model popularity.

Provider-declared conditions

fail closed

Catalog supply can carry data-use and access-review conditions into eligibility before pricing, demand, or performance scores run.

Variant Contract Lab

mode + meter

Resolve interactive, deferred, or deliberate supply with a distinct task contract, native meter, AI Token estimate, and capacity path.

Marketplace

Bring broad model choice into Aurona's routing layer.

The model supply is treated as a programmable catalog. Aurona differentiates through token credits, route scores, provider policy, private GPU lanes, and enterprise observability.

Showing 17 matched models. Click any model to update the route preview.

Sort

Top Models

Weekly model usage translated into Aurona route score signals

All · 6 plotted

Mon

Tue

Wed

Thu

Fri

Sat

Sun

Catalog

Representative model supply mapped into Aurona routes.

These rows use familiar model identifiers, pricing, context windows, and modality signals, then package them as production route candidates.

Source model

Aurona route

Context / modality

Input / output per 1M

openai/gpt-5.5-pro

Premium

1.1M / text+image+file

$30.00 / $180.00

anthropic/claude-opus-4.8

Quality

1M / text+image+file

$15.00 / $75.00

google/gemini-3.1-pro-preview

Multimodal

1.0M / text+image+file+audio+video

$1.50 / $9.00

deepseek/deepseek-v4-flash

Value flash

1.0M / text

$0.420 / $0.840

google/gemini-2.5-pro

Long context

1.0M / text+image+file+audio+video

$1.25 / $10.00

x-ai/grok-4.3

Fast reasoning

1M / text+image+file

$1.25 / $2.50

deepseek/deepseek-v4-pro

Value

1.0M / text

$0.435 / $0.870

qwen/qwen3.7-plus

Value

1M / text+image

$0.320 / $1.28

meta-llama/llama-4-maverick

Open weight

1.0M / text+image

$0.150 / $0.600

google/gemini-3-pro-image

Image output

66K / text+image->text+image

$2.00 / $12.00

nvidia/nemotron-3-ultra-550b-a55b

GPU reasoning

1M / text

$0.500 / $2.20

z-ai/glm-5.2

Value reasoning

1.0M / text

$0.950 / $3.00

aurona/sovereign-gpu

Sovereign

regional policy / private lane

Custom / Custom

aurona/app-builder-route

App-ready

marketplace meter / logs / tools

Metered / Metered

aurona/serverless-endpoint

Serverless

warm cache / queue policy / usage logs

Metered / Metered

aurona/dedicated-endpoint

Dedicated

pinned model / private lane / billing export

Custom / Custom

openai/gpt-oss-120b:free

Free onboarding

131K / text

Free / Free

Ranked routes

Choose by job, not by brand name.

The model page should help developers and enterprises pick the right route for coding, realtime UX, private workloads, and governed usage.

Job

Source model

Aurona signal

Route policy

Premium reasoning

openai/gpt-5.5-pro

98 route score

use when quality matters more than price

Agent quality

anthropic/claude-opus-4.8

96 route score

coding, tools, documents, enterprise assistants

Multimodal breadth

google/gemini-3.1-pro-preview

audio + video input

media-heavy apps and broad product coverage

Low-cost throughput

deepseek/deepseek-v4-flash

$0.420 input

automation with predictable token margins

Open-weight GPU lane

meta-llama/llama-4-maverick

$0.150 input

dedicated inference and private capacity

Image output

google/gemini-3-pro-image

text+image output

creative apps and generated media workflows

Sovereign capacity

aurona/sovereign-gpu

regional approval

regulated workloads with fail-closed routing

App marketplace

aurona/app-builder-route

app meter + logs

builders that need install, customer, and settlement attribution

Serverless endpoint

aurona/serverless-endpoint

queue + warm cache

bursty apps, evals, and developer endpoint trials

Dedicated endpoint

aurona/dedicated-endpoint

pinned model + private lane

enterprise tenants and latency-sensitive workloads

Free onboarding

openai/gpt-oss-120b:free

free route

developer trials, demos, and evaluation traffic

Catalog lifecycle

A production catalog should show what changes next.

Aurona can keep model discovery useful after launch by exposing maturity, version policy, regional fit, provider-declared output-use metadata, and a replacement route before retirement.

Signal

Catalog field

Operational use

Release state

preview, current, maintenance, retiring

Make maturity visible before a route is approved

Model age

published date and days in catalog

Separate a new launch from a stable production default

Version policy

floating alias or pinned version

Choose convenience for trials or reproducibility for production

Retirement plan

retirement date, replacement route, owner

Give apps and workspace owners a migration path

Regional availability

eligible regions and endpoint classes

Keep route selection aligned with approved capacity

Training review

provider-declared output-use metadata

Route dataset generation only after terms are reviewed

Stable alias

workspace-owned alias with an explicit update policy

Adopt newer model supply without silently changing pinned production routes

Catalog discovery

Filter the supply before the router scores it.

Aurona discovery can narrow models by output modality, request controls, provider-declared conditions, economics, performance, demand, and lifecycle before policy and capacity scoring begins.

Discovery lens

Catalog signal

Aurona use

Output modality

text, image, audio, embeddings

Keep one catalog and one credential across different generation workloads

Decision output

yes/no, choice, or score with a declared schema

Treat classification and review steps as typed workloads rather than chat text

Supported parameters

tools, structured output, reasoning, sampling

Filter out endpoints that cannot honor the request contract

Price

input, output, request, media unit

Compare the complete metering shape before estimating AI Token credits

Performance

latency and throughput bands

Shortlist healthy routes for interactive or batch workloads

Popularity

modeled weekly route demand

Use adoption as context, never as a substitute for workload evaluation

Catalog age

newest, current, maintenance, retiring

Separate launch discovery from production lifecycle policy

Supply conditions

provider-declared data use and access class

Require access and data-use review before a workspace alias can resolve

Catalog Query Lab · simulated

Query supply before a route becomes a decision.

Filter hard capabilities and provider-declared conditions, choose a measurable sort, resolve an alias, and inspect AI Token and GPU capacity evidence in one read-only receipt.

2 eligible · planning only

Identity policy

Query preview

/v1/models/query-preview?output=text&requires=tools&sort=throughput-high-to-low&capacity=all&data_use=no-contribution-only&access=standard-only

Receipt

catalog_01czphly

Mode

planning-only

Resolved identity · preview

aurona/catalog-choice

canonical → aurona/value-text-2026-07

#1 eligible

Value Text

aurona/value-text-2026-07

196 tokens/s p50

Lifecycle

current

Capacity

serverless

Est. AI Tokens

0.15

Declared conditions: none

#2 eligible

Agent Balanced

aurona/agent-balanced-2026-08

142 tokens/s p50

Lifecycle

current

Capacity

serverless

Est. AI Tokens

0.43

Declared conditions: none

Excluded before scoring

aurona/private-agent-2026-06 · access requires workspace review

Simulated provider-declared catalog evidence. This preview does not change aliases or policy, send traffic, debit AI Tokens, publish a model, or reserve GPU capacity.

Variant Contract Lab · simulated

Resolve how the model runs, not only which family wins.

Choose delivery and reasoning requirements, then preserve a distinct variant ID, native meter, task contract, modeled AI Token amount, and GPU capacity path.

Variant resolved · planning only

Variant query

/v1/models/variant-preview?family=aurona/atlas&delivery_mode=interactive&reasoning_class=standard

receipt · variant_00j6b6i4

Resolved variant · preview

aurona/atlas-2026-09

family → aurona/atlas

Task

synchronous

Native meter

input + output tokens

AI Tokens

3.8 modeled

Capacity

serverless

Next · send request and keep the variant receipt

Excluded before resolution

aurona/atlas-2026-09:batch · delivery requires a deferred job

aurona/atlas-2026-09:deliberate · reasoning class does not match

Simulated planning data. This preview does not send a request, create a job, promise latency or discounts, debit AI Tokens, change an alias, or reserve GPU capacity.

Protocol Route Lab · simulated

Keep the request contract visible through routing.

Preview native and reviewed-adapter paths, then verify an ordered fallback chain without silently changing tools, reasoning, streaming, structured output, or media job semantics.

fail-closed hold

/v1/routes/protocol-preview?protocol=responses&requires=tools&adapter=preserve-only&capacity=all&fallback=approved-chain&attempts=shared-contract

Receiptprotocol_01j305fi

No route preserves the requested contract.

No verified fallback preserves every protocol, feature, and capacity gate.

Fallback evidence: 0 verified attempts · shared-contract · triggers on provider unavailable, rate limit, or policy refusal.
Excluded evidence: aurona/private-messages-2026-08: protocol translation is not allowed · aurona/video-job-2026-08: tools would not be preserved

Planning only. This lab does not send traffic, execute a fallback, mutate a preset, debit AI Tokens, or reserve GPU capacity.

Decision lab · simulated

Turn catalog evidence into a route shortlist.

Compare modeled quality, intelligence, design fit, latency, cost, lifecycle, policy, and GPU readiness. Hard requirements remove incompatible supply before scoring.

4 eligible routes

Objective

#1 shortlist

Value Fast

aurona/auto-fast

fit 79

Quality

82

Intelligence

79

Design fit

72

Latency

280 ms

cost index 0.90currentshared GPU pool79 blended evidenceGPU lane ready

#2 shortlist

Open Private

aurona/private-gpu

fit 77

Quality

86

Intelligence

83

Design fit

78

Latency

430 ms

cost index 1.80currentreserved regional lane84 blended evidenceGPU lane ready

#3 shortlist

Agent Balanced

aurona/agent-balanced

fit 76

Quality

91

Intelligence

92

Design fit

84

Latency

520 ms

cost index 5.60currentserverless + reserved fallback90 blended evidenceGPU lane ready

Decision workloads

Give small decisions their own model contract.

A classification or score should declare its possible outcomes, evaluation method, native meter, and review threshold before it enters an Aurona route. This is a proposed planning pattern, not a live decision endpoint.

Planning field

Evidence

Aurona use

Task contract

Question type, allowed choices, score range, and output schema

Reject a route that cannot return the requested decision shape

Evaluation

Labeled cases, calibration, abstentions, and false-positive cost

Compare candidates on the application decision, not general model popularity

Meter

Declared request or token unit, then a modeled AI Token estimate

Keep decision economics separate from text-generation assumptions

Review path

Workspace policy, route trace, and human-review threshold

Make consequential outcomes inspectable before production use

Token economics

Every route needs visible commercial logic.

Transparent model access matters. Aurona goes further by showing token credits, provider health, settlement, and compute lanes.

Metric

How it is shown

Why it matters

Source model price

Displayed per 1M input and output tokens

Used as the benchmark for market clarity

Aurona token credit

Mapped to credits, prepaid balances, and enterprise invoices

Used for settlement and margin control

Route score

Quality + health + policy + cost

Used for default route selection

Provider health

Availability, latency, rate-limit, and moderation signals

Used for failover

GPU lane

Dedicated, partner, or spot capacity

Used for latency and throughput

Endpoint health

Queue depth, warm cache, quota, saturation, retry reason

Used for serverless and dedicated routing

Broadcast trace

Provider, fallback, cost, policy, and app attribution event

Used for logs, webhooks, and observability

Service tier

Shared, priority, reserved, dedicated, or sovereign

Used for capacity planning and procurement review

Workload evidence

A route should win on your application, not a generic leaderboard.

Aurona can turn saved prompt cases, tool assertions, quality thresholds, AI Token estimates, policy checks, and capacity forecasts into a repeatable route-promotion decision.

Workload evaluation · simulated

Promote a route with application evidence.

Pin one workload suite, compare route behavior on the same cases, and gate promotion by answer quality, tool assertions, workspace policy, AI Token use, and capacity fit.

3 routes pass

Evaluation suite

#1 · passes release gate

Frontier Quality

aurona/fusion-quality

eval 94

Quality

94

Cases

100%

Tools

100%

p95

920 ms

AI Tokens

32

keep in provider pool

#2 · passes release gate

Agent Balanced

aurona/agent-balanced

eval 91

Quality

88

Cases

100%

Tools

100%

p95

540 ms

AI Tokens

14

keep in serverless pool

#3 · passes release gate

Private GPU

aurona/private-gpu

eval 90

Quality

87

Cases

100%

Tools

100%

p95

470 ms

AI Tokens

17

keep in serverless pool

#4 · review required

Value Fast

aurona/auto-fast

eval 77

Quality

81

Cases

67%

Tools

67%

p95

310 ms

AI Tokens

7

keep in serverless poolmisses quality floormisses tool assertions

Demonstration data only. Scores, latency, credits, and capacity guidance are simulated; saving this preview does not change a production route.

Model operations

Everything around the model matters.

The winning platform is not only the model list. It is routing, policy, credits, observability, and compute access around the model.

Search and filter

Browse by model family, provider, capability, latency class, token price, context window, policy, and GPU-backed availability.

Compare routes

Compare model quality, input price, output price, context window, route score, fallback behavior, and enterprise fit.

Tune tradeoffs

Expose a cost-quality control so teams can choose cheapest healthy model, highest quality, or a balanced production route.

Provider policy

Decide which providers can receive traffic by app, customer, data sensitivity, region, and procurement status.

Credit forecasting

Estimate how much a model route will cost before a product team moves traffic into production.

Fallback design

Plan secondary providers and GPU lanes before rate limits, outages, or price changes affect users.

Key log handoff

Let a model test open the exact API key usage tab and pre-filtered logs that explain spend, retries, and policy blocks.

Endpoint diagnostics

Show queue depth, warm-cache state, quota, saturation, and retry reason before traffic moves to a serverless or dedicated lane.

Private supply

Attach proprietary models, fine-tuned models, and dedicated GPU-backed routes to the same catalog.

Aurona.ai

Build on the token network for AI applications, models, and compute.

API

OpenAI-compatible

Billing

AI Token credits

Capacity

GPU-backed routes