QuarluxAI Power Control

Same GPU cluster, same power budget: with RL vs without

WITH RL agent  183,069 tokens WITHOUT (throttled safe)  134,860 tokens power budget

What the agent buys you — measured

Without RL (throttled)With RL agentGain
Economics — output (tokens / 20 min)134,860183,069+35.7%
— cost per million tokens$0.082$0.065−22%
— carbon per million tokens257 g203 g−21%
Technical — energy per 1k tokens1.219 Wh0.947 Wh−22.3%
— throughput (tokens/s)114.5154.8+35%
— time over budget0%2.3%vs 17.7% uncontrolled
— power overshoot consistency159 ± 8 W·scontractable (demand response)

Simulate the agent on your facility

One input: your facility's power capacity in megawatts. Press simulate: watch the agent manage your fleet, then get your full year.
Your power capacity1 MW · ~1,430 GPUs
your fleet WITH agent WITHOUT (throttled) your budget
0
tokens WITH agent
0
tokens without
+0
extra tokens from the same electricity
electricity saved producing these tokens

Your fleet over one year,

Without RLWith RL agentYou gain
Your workload (same tokens as today) — electricity to produce it
— electricity cost / yr
— carbon / yr
— cost per million tokens
Bonus: run the freed capacity anyway — extra output available
— capacity you must provision
Technical — energy per 1k tokens1.219 Wh0.947 Wh−22.3%
— time over your budget0% (throttled) / 17.7% (uncontrolled)2.3%−87% vs uncontrolled
— overshoot consistency159±8 W·sdemand-response eligible

Plugging in: sidecar agent · reads GPU power, no root · sets your vLLM/TGI/Triton batch knob · one config value: your budget · auto-calibrates in 20 min · fail-safe reverts to your setting · week 1 read-only on your data

Get the 30-day pilot
Measured on live A100 hardware (3 replications, 2 machines): 38.7 vs 28.6 tokens/s per GPU, 131.9 vs 125.3 W per GPU, 0.947 vs 1.219 Wh per 1k tokens. Anchors: $0.08/kWh, 0.25 kg CO₂/kWh, 70% utilization, 700 W/GPU, $4M/MW — replaced by your data in the pilot. © QuarluxAI