Two published papers, and the benchmark results behind them. Real sites, real distance between them, no synthetic tests.
Research
Our claims are published, cited and available in full. Read the papers before you believe the marketing.
— Solyx AI
Where a request runs decides how much carbon it burns, because the grid is far dirtier in some places and at some hours than others. We showed live, on real GPUs across several regions, that inference can be steered toward the cleaner grid without a single failed request. Replaying a full year of grid data, that steering cuts the emissions of the same workload by roughly half.Read on arXiv →
, N. Katla — Solyx AI
Ten live signals — from the GPUs, the software and the network between sites — decide where each request runs. Measured against the standard approach on real hardware in three locations: about 1.6 to 1.75 times the useful work from the same fleet, a quarter off the slowest responses, and recovery from a failed site in about a second instead of four.Read on arXiv →
Head to head
Every figure below is measured, not modeled. Identical fleet, identical workload, identical SLO target — the only variable is the placement decision.
How it was measured
Two independent benchmark campaigns against a 216-cell SLO matrix, spanning the major inference workload classes. The control plane reads ten routing signals — four application-layer, four hardware, two network — and sits above unmodified inference engines, emitting standard Envoy xDS endpoint weights. The comparison baseline is a real production pressure-based router, not uniform placement.
Want to see the full methodology?
Let's Talk →