#1 eligible
Value Text
aurona/value-text-2026-07
Lifecycle
current
Capacity
serverless
Est. AI Tokens
0.15
Declared conditions: none
Models
Search public and private model supply, compare token economics, inspect route health, and choose the right provider or GPU-backed lane through Aurona.
OpenAI-compatible
Token ledger
GPU route layer
Source models
400+
Long context
1M+
Entry routes
Free
openai/gpt-5.5-pro
price, latency, policy, fallback
anthropic/claude-opus-4.8
price, latency, policy, fallback
deepseek/deepseek-v4-flash
price, latency, policy, fallback
sovereign/private
price, latency, policy, fallback
Model control plane
Aurona should treat each model as a productized supply object: price, context, modality, provider health, policy support, fallback depth, and compute path all matter.
Source catalog
600+ market models
Public, private, and partner model supply is mapped into Aurona routes with provider, context, modality, price, and endpoint signals.
Route score
0-100
Combines quality, provider health, policy fit, fallback depth, token price, and capacity readiness.
Policy gates
ZDR + region
Routes can be filtered by retention, approved vendors, parameter support, and customer tier.
Compute path
Hosted or GPU
High-volume routes can move from public provider APIs into private or reserved inference lanes.
Key mode
Aurona or BYOK
Route previews can distinguish Aurona-managed access, customer provider vault keys, and hybrid fallback.
Service tier
5 classes
Shared, priority, serverless, reserved, and dedicated routes help customers match latency and capacity to workload risk.
Endpoint health
queue + cache
Dedicated and serverless routes can expose queue depth, warm-cache state, quota, saturation, and retry reason.
Policy actions
block/mask/warn
Catalog tests can preview which route, provider, prompt, output, or budget rules would fire before production traffic is sent.
Prompt history
saved runs
Approved playground tests can carry model parameters, modality, route score, credit estimate, and redaction state into launch review.
Trace fanout
broadcast
Route traces can be copied into usage logs, webhooks, observability tools, and billing exports.
Key detail
usage + logs
Every catalog test can link back to the API key, budget progress, request filters, and route history.
Admin API
catalog to control
Model rows can become route aliases, budget rules, provider-key policies, and app meters through management endpoints.
Regional fit
sovereign
Model rows can expose region, vendor, and capacity constraints before a regulated route is approved.
Catalog lifecycle
launch to retire
Model age, preview state, version pinning, replacement route, and retirement date can be reviewed before an app depends on a slug.
Output use
training review
Provider-declared training or distillation eligibility can be surfaced as review metadata instead of being inferred from model popularity.
Provider-declared conditions
fail closed
Catalog supply can carry data-use and access-review conditions into eligibility before pricing, demand, or performance scores run.
Variant Contract Lab
mode + meter
Resolve interactive, deferred, or deliberate supply with a distinct task contract, native meter, AI Token estimate, and capacity path.
Marketplace
The model supply is treated as a programmable catalog. Aurona differentiates through token credits, route scores, provider policy, private GPU lanes, and enterprise observability.
Showing 17 matched models. Click any model to update the route preview.
Weekly model usage translated into Aurona route score signals
Mon
Tue
Wed
Thu
Fri
Sat
Sun
Catalog
These rows use familiar model identifiers, pricing, context windows, and modality signals, then package them as production route candidates.
Source model
Aurona route
Context / modality
Input / output per 1M
openai/gpt-5.5-pro
Premium
1.1M / text+image+file
$30.00 / $180.00
anthropic/claude-opus-4.8
Quality
1M / text+image+file
$15.00 / $75.00
google/gemini-3.1-pro-preview
Multimodal
1.0M / text+image+file+audio+video
$1.50 / $9.00
deepseek/deepseek-v4-flash
Value flash
1.0M / text
$0.420 / $0.840
google/gemini-2.5-pro
Long context
1.0M / text+image+file+audio+video
$1.25 / $10.00
x-ai/grok-4.3
Fast reasoning
1M / text+image+file
$1.25 / $2.50
deepseek/deepseek-v4-pro
Value
1.0M / text
$0.435 / $0.870
qwen/qwen3.7-plus
Value
1M / text+image
$0.320 / $1.28
meta-llama/llama-4-maverick
Open weight
1.0M / text+image
$0.150 / $0.600
google/gemini-3-pro-image
Image output
66K / text+image->text+image
$2.00 / $12.00
nvidia/nemotron-3-ultra-550b-a55b
GPU reasoning
1M / text
$0.500 / $2.20
z-ai/glm-5.2
Value reasoning
1.0M / text
$0.950 / $3.00
aurona/sovereign-gpu
Sovereign
regional policy / private lane
Custom / Custom
aurona/app-builder-route
App-ready
marketplace meter / logs / tools
Metered / Metered
aurona/serverless-endpoint
Serverless
warm cache / queue policy / usage logs
Metered / Metered
aurona/dedicated-endpoint
Dedicated
pinned model / private lane / billing export
Custom / Custom
openai/gpt-oss-120b:free
Free onboarding
131K / text
Free / Free
Ranked routes
The model page should help developers and enterprises pick the right route for coding, realtime UX, private workloads, and governed usage.
Job
Source model
Aurona signal
Route policy
Premium reasoning
openai/gpt-5.5-pro
98 route score
use when quality matters more than price
Agent quality
anthropic/claude-opus-4.8
96 route score
coding, tools, documents, enterprise assistants
Multimodal breadth
google/gemini-3.1-pro-preview
audio + video input
media-heavy apps and broad product coverage
Low-cost throughput
deepseek/deepseek-v4-flash
$0.420 input
automation with predictable token margins
Open-weight GPU lane
meta-llama/llama-4-maverick
$0.150 input
dedicated inference and private capacity
Image output
google/gemini-3-pro-image
text+image output
creative apps and generated media workflows
Sovereign capacity
aurona/sovereign-gpu
regional approval
regulated workloads with fail-closed routing
App marketplace
aurona/app-builder-route
app meter + logs
builders that need install, customer, and settlement attribution
Serverless endpoint
aurona/serverless-endpoint
queue + warm cache
bursty apps, evals, and developer endpoint trials
Dedicated endpoint
aurona/dedicated-endpoint
pinned model + private lane
enterprise tenants and latency-sensitive workloads
Free onboarding
openai/gpt-oss-120b:free
free route
developer trials, demos, and evaluation traffic
Catalog lifecycle
Aurona can keep model discovery useful after launch by exposing maturity, version policy, regional fit, provider-declared output-use metadata, and a replacement route before retirement.
Signal
Catalog field
Operational use
Release state
preview, current, maintenance, retiring
Make maturity visible before a route is approved
Model age
published date and days in catalog
Separate a new launch from a stable production default
Version policy
floating alias or pinned version
Choose convenience for trials or reproducibility for production
Retirement plan
retirement date, replacement route, owner
Give apps and workspace owners a migration path
Regional availability
eligible regions and endpoint classes
Keep route selection aligned with approved capacity
Training review
provider-declared output-use metadata
Route dataset generation only after terms are reviewed
Stable alias
workspace-owned alias with an explicit update policy
Adopt newer model supply without silently changing pinned production routes
Catalog discovery
Aurona discovery can narrow models by output modality, request controls, provider-declared conditions, economics, performance, demand, and lifecycle before policy and capacity scoring begins.
Discovery lens
Catalog signal
Aurona use
Output modality
text, image, audio, embeddings
Keep one catalog and one credential across different generation workloads
Decision output
yes/no, choice, or score with a declared schema
Treat classification and review steps as typed workloads rather than chat text
Supported parameters
tools, structured output, reasoning, sampling
Filter out endpoints that cannot honor the request contract
Price
input, output, request, media unit
Compare the complete metering shape before estimating AI Token credits
Performance
latency and throughput bands
Shortlist healthy routes for interactive or batch workloads
Popularity
modeled weekly route demand
Use adoption as context, never as a substitute for workload evaluation
Catalog age
newest, current, maintenance, retiring
Separate launch discovery from production lifecycle policy
Supply conditions
provider-declared data use and access class
Require access and data-use review before a workspace alias can resolve
Catalog Query Lab · simulated
Filter hard capabilities and provider-declared conditions, choose a measurable sort, resolve an alias, and inspect AI Token and GPU capacity evidence in one read-only receipt.
Identity policy
Query preview
/v1/models/query-preview?output=text&requires=tools&sort=throughput-high-to-low&capacity=all&data_use=no-contribution-only&access=standard-only
Receipt
catalog_01czphly
Mode
planning-only
Resolved identity · preview
aurona/catalog-choice
canonical → aurona/value-text-2026-07
#1 eligible
aurona/value-text-2026-07
Lifecycle
current
Capacity
serverless
Est. AI Tokens
0.15
Declared conditions: none
#2 eligible
aurona/agent-balanced-2026-08
Lifecycle
current
Capacity
serverless
Est. AI Tokens
0.43
Declared conditions: none
Excluded before scoring
aurona/private-agent-2026-06 · access requires workspace review
Simulated provider-declared catalog evidence. This preview does not change aliases or policy, send traffic, debit AI Tokens, publish a model, or reserve GPU capacity.
Variant Contract Lab · simulated
Choose delivery and reasoning requirements, then preserve a distinct variant ID, native meter, task contract, modeled AI Token amount, and GPU capacity path.
Variant query
/v1/models/variant-preview?family=aurona/atlas&delivery_mode=interactive&reasoning_class=standard
receipt · variant_00j6b6i4
Resolved variant · preview
family → aurona/atlas
Task
synchronous
Native meter
input + output tokens
AI Tokens
3.8 modeled
Capacity
serverless
Next · send request and keep the variant receipt
Excluded before resolution
aurona/atlas-2026-09:batch · delivery requires a deferred job
aurona/atlas-2026-09:deliberate · reasoning class does not match
Simulated planning data. This preview does not send a request, create a job, promise latency or discounts, debit AI Tokens, change an alias, or reserve GPU capacity.
Protocol Route Lab · simulated
Preview native and reviewed-adapter paths, then verify an ordered fallback chain without silently changing tools, reasoning, streaming, structured output, or media job semantics.
/v1/routes/protocol-preview?protocol=responses&requires=tools&adapter=preserve-only&capacity=all&fallback=approved-chain&attempts=shared-contract
No route preserves the requested contract.
No verified fallback preserves every protocol, feature, and capacity gate.
Planning only. This lab does not send traffic, execute a fallback, mutate a preset, debit AI Tokens, or reserve GPU capacity.
Decision lab · simulated
Compare modeled quality, intelligence, design fit, latency, cost, lifecycle, policy, and GPU readiness. Hard requirements remove incompatible supply before scoring.
Objective
#1 shortlist
aurona/auto-fast
Quality
82
Intelligence
79
Design fit
72
Latency
280 ms
#2 shortlist
aurona/private-gpu
Quality
86
Intelligence
83
Design fit
78
Latency
430 ms
#3 shortlist
aurona/agent-balanced
Quality
91
Intelligence
92
Design fit
84
Latency
520 ms
Decision workloads
A classification or score should declare its possible outcomes, evaluation method, native meter, and review threshold before it enters an Aurona route. This is a proposed planning pattern, not a live decision endpoint.
Planning field
Evidence
Aurona use
Task contract
Question type, allowed choices, score range, and output schema
Reject a route that cannot return the requested decision shape
Evaluation
Labeled cases, calibration, abstentions, and false-positive cost
Compare candidates on the application decision, not general model popularity
Meter
Declared request or token unit, then a modeled AI Token estimate
Keep decision economics separate from text-generation assumptions
Review path
Workspace policy, route trace, and human-review threshold
Make consequential outcomes inspectable before production use
Token economics
Transparent model access matters. Aurona goes further by showing token credits, provider health, settlement, and compute lanes.
Metric
How it is shown
Why it matters
Source model price
Displayed per 1M input and output tokens
Used as the benchmark for market clarity
Aurona token credit
Mapped to credits, prepaid balances, and enterprise invoices
Used for settlement and margin control
Route score
Quality + health + policy + cost
Used for default route selection
Provider health
Availability, latency, rate-limit, and moderation signals
Used for failover
GPU lane
Dedicated, partner, or spot capacity
Used for latency and throughput
Endpoint health
Queue depth, warm cache, quota, saturation, retry reason
Used for serverless and dedicated routing
Broadcast trace
Provider, fallback, cost, policy, and app attribution event
Used for logs, webhooks, and observability
Service tier
Shared, priority, reserved, dedicated, or sovereign
Used for capacity planning and procurement review
Workload evidence
Aurona can turn saved prompt cases, tool assertions, quality thresholds, AI Token estimates, policy checks, and capacity forecasts into a repeatable route-promotion decision.
Workload evaluation · simulated
Pin one workload suite, compare route behavior on the same cases, and gate promotion by answer quality, tool assertions, workspace policy, AI Token use, and capacity fit.
Evaluation suite
#1 · passes release gate
aurona/fusion-quality
Quality
94
Cases
100%
Tools
100%
p95
920 ms
AI Tokens
32
#2 · passes release gate
aurona/agent-balanced
Quality
88
Cases
100%
Tools
100%
p95
540 ms
AI Tokens
14
#3 · passes release gate
aurona/private-gpu
Quality
87
Cases
100%
Tools
100%
p95
470 ms
AI Tokens
17
#4 · review required
aurona/auto-fast
Quality
81
Cases
67%
Tools
67%
p95
310 ms
AI Tokens
7
Demonstration data only. Scores, latency, credits, and capacity guidance are simulated; saving this preview does not change a production route.
Model operations
The winning platform is not only the model list. It is routing, policy, credits, observability, and compute access around the model.
Browse by model family, provider, capability, latency class, token price, context window, policy, and GPU-backed availability.
Compare model quality, input price, output price, context window, route score, fallback behavior, and enterprise fit.
Expose a cost-quality control so teams can choose cheapest healthy model, highest quality, or a balanced production route.
Decide which providers can receive traffic by app, customer, data sensitivity, region, and procurement status.
Estimate how much a model route will cost before a product team moves traffic into production.
Plan secondary providers and GPU lanes before rate limits, outages, or price changes affect users.
Let a model test open the exact API key usage tab and pre-filtered logs that explain spend, retries, and policy blocks.
Show queue depth, warm-cache state, quota, saturation, and retry reason before traffic moves to a serverless or dedicated lane.
Attach proprietary models, fine-tuned models, and dedicated GPU-backed routes to the same catalog.
Aurona.ai
API
OpenAI-compatible
Billing
AI Token credits
Capacity
GPU-backed routes