Frontier Research
quality firstGPT-5.5 Pro, Claude Opus 4.8, Gemini 3.1 Pro Preview
Premium reasoning judge
Best for source-heavy analysis, executive decisions, strategy memos, and answers where disagreement matters.
Fusion
Route one request through a panel of models, compare their reasoning, synthesize the best answer, enforce provider and data policy, and settle the full cost through Aurona credits.
OpenAI-compatible
Token ledger
GPU route layer
aurona/fusion
multi-model route
Panel models
3-7
Judge output
Structured
Fallback
Automatic
Classify
Detect task type, risk level, required modality, region, and budget before any model call.
Panel
Run approved frontier, value, open-weight, or private GPU models in parallel.
Judge
Compare consensus, contradictions, evidence gaps, unique insights, and confidence.
Synthesize
Return one final answer, route trace, token cost, provider health, and fallback notes.
Fusion control layer
Fusion should expose the knobs developers actually need: panel membership, judge behavior, provider routing, credit ceiling, fallback policy, and trace output.
Panel control
3-7 models
Choose frontier, value, code, multimodal, or private GPU models for each Fusion run.
Judge schema
Structured
Return consensus, contradictions, confidence, coverage gaps, unique insights, and final recommendation.
Provider policy
sort + order
Tune provider selection by price, throughput, latency, fallback permissions, and data collection policy.
Endpoint policy
queue + cache
Route panels through provider pools, serverless endpoints, or dedicated lanes using queue depth, warm-cache, and quota signals.
Cache affinity
traceable reuse
Keep repeated system prompts, tool schemas, and agent sessions on a compatible warm route when policy and health still match.
Credit guardrail
budget ceiling
Set max tokens, max credit, trace depth, and fail-closed rules before a panel starts.
Route modes
Model choice and provider routing are only the base layer. Aurona Fusion turns them into a full decision router with panel models, judge models, private supply, and token settlement.
Route
Primary signal
Best for
Execution
aurona/fusion
multi-model deliberation
research, agents, high-risk answers
panel + judge + final
aurona/fusion-fast
small panel, low latency
chat, copilots, realtime UX
2 panel models + fast judge
aurona/fusion-code
coding model panel
repository tasks, PR review, migration
code score + tool use
aurona/fusion-private
approved private supply
enterprise data, regulated workloads
ZDR + region + GPU lane
aurona/fusion-endpoint
serverless or dedicated endpoint panel
app bursts, evals, private tenants
endpoint health + judge trace
aurona/auto
quality, price, latency, health
default production traffic
single best route
aurona/exact
pinned model and provider
stable behavior, audits, certification
no automatic substitution
Fusion simulator
Choose a route preset and run it. The panel, chart, and trace update immediately.
quality
96
cost control
68
latency
74
Trace #18
aurona/fusion ran a panel, checked zero-retention policy, wrote token ledger events, and returned a judge confidence score of 96.
classify
task, risk, modality, budget
panel
4 approved model calls
judge
96% confidence
ledger
1.73 credit ceiling
{
"model": "aurona/fusion",
"panel_size": 4,
"judge": "aurona/judge-premium",
"provider": {
"sort": "throughput",
"allow_fallbacks": true,
"cost_quality_tradeoff": 7
},
"endpoint": {
"type": "serverless",
"health_policy": ["queue_depth", "warm_cache", "saturation"]
},
"policy": "zero_retention_optional",
"session_id": "fusion-preview-18",
"observability": { "destination": "route-trace-stream" },
"budget": { "max_credit": "1.73" }
}Auto Route decision lab
simulated planning dataClassify the workload, score only approved routes, choose a cost-quality objective, and return a compact receipt with fallback and capacity evidence.
Objective
Allowed pool
Decision receipt #2048
aurona/code-balanced
fallback · aurona/private-code
workload
Code
route score
109
credit estimate
0.048
capacity
Shared + burst GPU
Code workload detected before generation.
Balanced objective scored quality, latency, and simulated credit exposure.
5 routes remained after the all approved pool check.
Shared + burst GPU matched the selected route's current planning profile.
{
"simulated": true,
"objective": "Balanced",
"scope": "All approved",
"task": "Code",
"poolSize": 5,
"selectedRoute": "aurona/code-balanced",
"fallbackRoute": "aurona/private-code",
"capacity": "Shared + burst GPU",
"routeScore": 109,
"estimatedCredits": "0.048"
}Session continuity lab
simulated planning dataPreview whether a multi-turn session can reuse its route, should switch models or GPU lanes, or must pause for policy and AI Token review. This lab does not pin production traffic.
Next task
Route health
Workspace policy
AI Token gate
GPU capacity preference
Continuity receipt #731
reuse route
aurona/code-balanced
prior route
aurona/code-balanced
next task
Code
capacity
Serverless approved pool
Code remained compatible with the prior session route.
Route health remained eligible for continuity planning.
Workspace policy remained unchanged and eligible.
The AI Token budget gate remained open.
Tool Route Evidence Lab
simulated · planning onlyCheck the required tool schema and capacity scope first, then compare simulated validity, evaluation, throughput, and AI Token evidence with a visible fallback.
Preview #417
No traffic, billing, or capacity change
{
"route": "aurona/agent-balanced",
"workload": "Support actions",
"tools": {
"schema_contract": "strict",
"parallel": true
},
"objective": "balanced",
"capacity_scope": "workspace-approved",
"mode": "preview"
}Selected endpoint
aurona/endpoint-cobalt
Reviewed fallback
aurona/endpoint-cedar
| Endpoint | Score | Schema valid | Eval confidence | Throughput | AI Tokens |
|---|---|---|---|---|---|
| aurona/endpoint-cobalt | 91.2 | 99% | 95% | 92 | 0.038 |
| aurona/endpoint-cedar | 87.9 | 99% | 92% | 76 | 0.052 |
| aurona/endpoint-onyx | 87.8 | 98% | 90% | 88 | 0.044 |
| aurona/endpoint-amber | 86.9 | 97% | 84% | 112 | 0.031 |
strict tool schema and parallel execution checked
4 approved endpoints remained after capability and capacity policy
balanced evidence ranked schema validity, evaluation confidence, throughput, and AI Token exposure
aurona/endpoint-cobalt selected with aurona/endpoint-cedar retained as the reviewed fallback
Panel presets
Each preset defines which models join the panel, which judge evaluates them, what budget limits apply, and how much route trace should be returned.
GPT-5.5 Pro, Claude Opus 4.8, Gemini 3.1 Pro Preview
Premium reasoning judge
Best for source-heavy analysis, executive decisions, strategy memos, and answers where disagreement matters.
Claude Opus 4.8, Kimi K2.7 Code, DeepSeek R1, Qwen3.7 Plus
Code-aware judge
Best for complex coding, migrations, debugging, architecture review, and tool-heavy developer agents.
GPT-5.5 Pro, Grok 4.3, Gemini 3.1 Pro Preview
Evidence judge
Best for fresh market reads, competitive analysis, policy monitoring, and multi-source synthesis.
Llama 4 Maverick, Nemotron Ultra, DeepSeek V4 Pro
Private enterprise judge
Best for customers that need controlled data paths, predictable throughput, and dedicated inference capacity.
DeepSeek V4 Flash, Llama 4 Maverick, Qwen3.7 Plus
Endpoint-aware judge
Best for apps that need burstable serverless capacity first, then a dedicated lane after traffic becomes predictable.
API configuration
A Fusion request should expose the important controls directly: panel models, judge behavior, fallback policy, provider sorting, zero-retention requirements, and credit budget.
Fusion request
{
"model": "aurona/fusion",
"messages": [{ "role": "user", "content": "Evaluate this acquisition memo." }],
"fusion": {
"panel": ["openai/gpt-5.5-pro", "anthropic/claude-sonnet-4.6", "google/gemini-3.5-flash"],
"judge": "aurona/judge-premium",
"return_analysis": true,
"budget": { "max_tokens": 180000, "max_credit": "12.00" }
},
"provider": {
"sort": "throughput",
"allow_fallbacks": true,
"cost_quality_tradeoff": 7,
"cache_affinity": "best_effort",
"cache_key": "acquisition-memo-tools-v2",
"data_collection": "deny",
"require_parameters": true
},
"endpoint": {
"type": "serverless",
"health_policy": ["queue_depth", "warm_cache", "quota"]
},
"session_id": "acquisition-memo-review",
"observability": { "destination": "workspace-usage-warehouse" }
}Choose a fixed panel, a cost-aware panel, or an adaptive panel based on task risk, context size, modality, and customer tier.
Sort by price, latency, throughput, or provider order while respecting region, retention, parameter support, and vendor approval.
Define exactly when Fusion may retry, downgrade, switch providers, use private GPU capacity, or fail closed for regulated traffic.
Ask the judge to return consensus, contradictions, confidence, evidence gaps, policy flags, and recommended next actions.
Track model tokens, judge tokens, final tokens, orchestration fees, and customer credits in one route ledger.
Return model IDs, providers, route decisions, fallback reasons, policy checks, and cost events for audit and procurement teams.
Judge analysis
Instead of hiding the panel, Aurona can expose a structured analysis object that lets developers inspect quality, disagreement, confidence, and cost.
Field
What it means
Product use
Consensus
Where panel models agree
Returned as the safest shared answer
Contradictions
Where models disagree
Flagged before final synthesis
Coverage gaps
What the panel missed
Used to trigger retry or retrieval
Unique insights
Distinct high-value observations
Merged into the final answer
Confidence
Evidence, model agreement, and policy fit
Shown in the route trace
Cost trace
Input, output, judge, and final tokens
Settled through Aurona credits
Provider controls
Fusion should support provider sorting, fallbacks, price ceilings, parameter requirements, and data collection controls while also adding Aurona's region, vendor approval, and GPU capacity logic.
Control
Example value
Why it matters
sort
price, throughput, latency
Choose the cheapest, fastest, or highest-throughput healthy provider
order
approved provider list
Prefer enterprise-contracted providers before public fallback
allow_fallbacks
true or false
Control whether Fusion can substitute when a provider fails
cost_quality_tradeoff
0-10
Tune the route between cost control and answer quality
session_id
string
Keep related agent turns pinned to a compatible provider when cache and policy allow
cache_affinity
off, best_effort, required
Prefer a compatible warm provider without bypassing policy, price, or health gates
cache_key
string
Group reusable system prompts, tool schemas, and agent turns into one traceable cache scope
max_price
per-token ceiling
Block accidental premium spend on large or repeated calls
data_collection
allow or deny
Respect zero-retention and sensitive-data policy
endpoint.type
provider_pool, serverless, dedicated
Choose which capacity class can run a panel member
endpoint.health_policy
queue_depth, warm_cache, quota
Avoid cold or saturated endpoints during panel execution
require_parameters
true or false
Only route to providers that support required parameters
Token economics
A multi-model route creates a new commercial object: panel cost, judge cost, final synthesis, orchestration fee, and GPU capacity can all be priced, settled, and optimized.
Cost layer
Measured from
Business value
Panel calls
Input and output tokens for each selected model
Makes multi-model quality measurable
Cached input
Provider-reported cache reads and writes
Separates reusable context from fresh prompt input in the credit trace
Judge pass
Structured comparison tokens
Turns disagreement into product signal
Final synthesis
One final model call or deterministic merge
Keeps the user experience simple
Fusion fee
Aurona route orchestration and settlement
Creates the AI Token revenue layer
GPU lane
Dedicated or partner inference capacity
Improves margin once volume is predictable
Aurona.ai
API
OpenAI-compatible
Billing
AI Token credits
Capacity
GPU-backed routes