Benchmarks/OwnX
01 · OwnX · Across the stack

Ownership,
not rental.

Eight benchmarks on the thing nobody else measures: what you still have when the vendor changes its mind.

$15,388Saved over three yearsFive-person team, modelled
4 of 4Things that survivePrice rise · outage · shutdown · cancel
5.0Ownership indexOur framework, of 5
0Token metersNo metering anywhere in the stack
8 benchmarks · 0 measured on hardware · 5 measure ownership
01

Own the stack instead of renting it: $15,388 stays with you over three years.

Cumulative cost of the same five-person workload over 36 months.

ModelledOwnershipModelled. Arithmetic on stated inputs. The inputs are named on the card.Method 36 months, five people. Hardware amortised at zero residual value; resold at 30% the gap widens by $720. Inputs [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$24,000$18,000$12,000$6,000$0mo 0mo 8mo 16mo 24mo 32mo 36break-even · month 8.4Rent$19,800Own$4,412cumulative cost · five-person teamRent = $110/mo stack at published list. Own = $2,400 hardware plus measured electricity at $0.15/kWh.
$15,388Gap at month 36
4.5×Rent multiple
8.4 moBreak-even

Method. 36 months, five people. Hardware amortised at zero residual value; resold at 30% the gap widens by $720. Inputs [C].

02

Price rise, outage, shutdown, ban — OwnX survives all five. Hosted AI survives none.

Whether the workload still runs and the history is still yours, under five failure modes.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Five failure modes, five stack types. 'Survives' means the workload still runs and the history is still yours. Our definition, published.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Price riseOutageShutdownYou cancelRegion blockscoreOwnX5/5Hosted AI subscription0/5Managed cloud GPU0.5/5Open weights, their host1/5Local-only tool5/5No vendor is named because the answer is the same for all of them.
5 of 5OwnX
0 of 5Hosted AI
5Failure modes tested

Method. Five failure modes, five stack types. 'Survives' means the workload still runs and the history is still yours. Our definition, published.

03

Ownership, scored: OwnX 5.0. The best alternative manages 3.8.

Our ownership framework applied to five stack types.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Our criteria, scored by us. This is not a measurement and is not presented as one.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
ControlPortabilityOptionalityPrivacyCostavgOwnX555555.0Hosted AI112121.4Managed GPU212111.4Open weights, hosted343233.0Local-only tools543223.2Scoring rules are on the method page so they can be argued with.
5.0OwnX average
1.3Hosted AI
3.8×Gap

Method. Our criteria, scored by us. This is not a measurement and is not presented as one.

04

Five vendors quietly cut what your money buys. Our terms moved 0%.

Change in what five vendors gave you, against the terms in force eighteen months earlier.

From published sourcesCandourFrom published sources. Read from public pricing pages, contracts or documentation.Method Our own accounts and invoices, Feb 2025 – Aug 2026. Vendors anonymised pending legal review; dated screenshots are in the data room [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0%30%60%90%120%Credits per dollar, vendor A-90%Images per month, vendor B-95%Rate limit, vendor C-62%Exportable context, vendor D-40%Price of the same tier, vendor E110%OwnX terms0%Negative is a reduction you received without asking for it.
5 of 5Vendors that changed terms
−67%Median change
0Changes to OwnX terms

Method. Our own accounts and invoices, Feb 2025 – Aug 2026. Vendors anonymised pending legal review; dated screenshots are in the data room [C].

05

Nobody else owns all five layers — price index, cloud, models, OS, device.

Which layers of the stack each kind of vendor actually owns.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Layer definitions are ours and published. Category comparison, not a vendor comparison.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Price indexCloudModelsOSDevicescoreOwnX5/5Hyperscaler1.5/5Frontier AI lab1/5GPU marketplace1.5/5Consumer AI box1.5/5Open-source stack2/5The point is the shape of the row, not any individual cell.
5 of 5OwnX layers
2.5 of 5Best competitor
6Categories

Method. Layer definitions are ours and published. Category comparison, not a vendor comparison.

06

Leaving OwnX takes half a day. Leaving an enterprise contract takes 210.

Days from the decision to leave until the same workload runs somewhere else.

From published sourcesOwnershipFrom published sources. Read from public pricing pages, contracts or documentation.Method Timed on our own accounts where an exit was available; otherwise read from standard contract terms [C]. 18 exit attempts in total.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 days150 days300 days450 days600 dayslow · median · highOwnX0 daysGPU marketplace12 daysHosted AI, paid18 daysManaged cloud74 daysColocation120 daysEnterprise contract210 daysRange across three exit attempts per vendor category.
0.5 dOwnX worst case
210 dEnterprise median
18Exits timed

Method. Timed on our own accounts where an exit was available; otherwise read from standard contract terms [C]. 18 exit attempts in total.

07

Consumer AI got 60% dearer in a year. Our price didn't move.

Cumulative change in published list price by vendor category, against Q3 2025.

From published sourcesCostFrom published sources. Read from public pricing pages, contracts or documentation.Method Published list prices captured quarterly from public pricing pages [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
shaded within each quarter · Q3 2025 = 0Q3 25Q4 25Q1 26Q2 26Q3 26Hosted AI, consumer0%15%15%40%40%Hosted AI, pro0%0%25%25%60%Managed GPU0%-5%-5%10%10%Frontier API0%-20%-20%-35%-35%GPU marketplace0%-12%-24%-31%-38%OwnX0%0%0%0%0%Consumer subscription tiers rise while raw API and marketplace prices fall.
+60%Worst consumer rise
−38%Best marketplace fall
0%OwnX

Method. Published list prices captured quarterly from public pricing pages [C].

08

Renting costs $19,800 over three years. Owning costs $4,412.

Where three years of money goes, costed line by line for both stacks.

ModelledCostModelled. Arithmetic on stated inputs. The inputs are named on the card.Method Overage taken from our own invoices as a proportion of base. Electricity measured at 75 W sustained, $0.15/kWh, 40% duty cycle.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0$6,000$12,000$18,000$24,000$13,200$3,960$2,640Subscriptions $13,200Overage $3,960Egress $2,640Rent, 36 months$19,800$2,400Hardware $2,400Electricity $1,012Replacement fund $1,000Own, 36 months$4,412Same five-person workload, same 36 months.
$19,800Rent total
$4,412Own total
78%Less

Method. Overage taken from our own invoices as a proportion of base. Electricity measured at 75 W sustained, $0.15/kWh, 40% duty cycle.

02 · Earth Compute · Live

Compute has
a price now.

Oil has a price. Gold has a price. Compute didn't — so we built the index, made it public, and kept the history.

44Provider feeds tracked27 pull live from the provider's own API
120,389SKUs trackedBare metal and cloud, all feeds
87Countries pricedServers with a price attached
6,621Facilities geolocatedPeeringDB + OSM, deduplicated
2,837Priced cellsEach linked to its source evidence
62%Global coverageAgainst a 90% target
18 benchmarks · 0 measured on hardware · 4 measure ownership
01

The same GPU server costs 5.2× more if you don't shop. We shop 34 providers every hour.

The same 8×H100 configuration, priced across every provider we index, in one hour.

From the live indexCostFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Live index snapshot, on-demand list pricing, spot excluded, captured within a single hour on 12 Aug 2026 [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0.00/hr$2.50/hr$5.00/hr$7.50/hr$10.00/hrlow · median · highCheapest tier$1.72/hrSecond tier$2.10/hrMarket median$4.35/hrWell-known provider$6.80/hrHyperscaler list$8.94/hrRange is across all indexed providers in that tier, same hour.
5.2×Cheapest to dearest
$7.22Spread per hour
34Providers

Method. Live index snapshot, on-demand list pricing, spot excluded, captured within a single hour on 12 Aug 2026 [A].

02

We can see 62% of the world's compute. The target is 90%, and the gap is published.

Share of identified global compute capacity present in the index.

From the live indexCandourFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Indexed capacity over best available estimate of total addressable capacity [A]. An estimate over an estimate, and labelled as one.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0%25%50%75%100%target 90%Global compute coverage62%+28 pt to goThe denominator is itself an estimate, which is why the gap is published rather than a claim of completeness.
62%Coverage today
90%Target
28 ptDistance to target

Method. Indexed capacity over best available estimate of total addressable capacity [A]. An estimate over an estimate, and labelled as one.

03

6,621 data centres mapped — 31× more than the biggest number anyone else has.

Data-centre facilities counted, against the only two public numbers that exist.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Ours: PeeringDB + OpenStreetMap facilities, deduplicated by location, 24 Aug 2026 [A]. Google: ABI Research 2025 campus estimate. AWS: leaked 2023 data.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 facilities2,000 facilities4,000 facilities6,000 facilities8,000 facilitiesEarth Compute, geolocated + deduplicated6,621 facilitiesAWS, leaked count215 facilities-97%Google, estimated campuses42 facilities-99%Published by anyone else0 facilities-100%Counting definitions differ — directional, not like for like.
6,621Facilities
31×vs leaked AWS count
3Counting methods

Method. Ours: PeeringDB + OpenStreetMap facilities, deduplicated by location, 24 Aug 2026 [A]. Google: ABI Research 2025 campus estimate. AWS: leaked 2023 data.

04

Beyond servers: 10,674 fibre routes, 1,326 exchange points, 12 gigawatts of power.

The four physical layers indexed alongside the machines.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Snapshot of 24 Aug 2026 [A]. Each layer carries its own open licence — OpenStreetMap ODbL, PeeringDB CC-BY, Epoch AI CC-BY — and its own refresh cadence.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
log scaleServer SKUs tracked114,600Fibre routes, terrestrial10,674Exchange points1,326Intercontinental links632Log scale. Fibre from OpenStreetMap, exchange points from PeeringDB, subsea and cross-border links counted separately. Plus 12 GW of IT power.
10,674Fibre routes
1,326Exchange points
12 GWIT power capacity

Method. Snapshot of 24 Aug 2026 [A]. Each layer carries its own open licence — OpenStreetMap ODbL, PeeringDB CC-BY, Epoch AI CC-BY — and its own refresh cadence.

05

Twenty-one live catalogues in one query — from Azure's 70,276 products to Hostkey's 42.

Live inventory carried per provider — 900× between the largest and the smallest.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Live feed table, 24 Aug 2026 [A]. Counting conventions differ by provider; each is counted as it publishes itself, so these compare catalogue size, not capacity.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
log scaleMicrosoft Azure70,276 productsAmazon Web Services21,901 productsGoogle Cloud15,983 productsVast.ai3,926 productsVultr3,517 productsHivelocity2,392 productsPhoenixNap456 productsRunPod338 productsScaleway160 productsDataCrunch143 productsHyperstack87 productsOracle Cloud81 productsHetzner76 productsFasthosts76 productsLinode (Akamai)75 productsOpenMetal72 productsMassed Compute69 productsLeafCloud56 productsCherry Servers49 productsOVHcloud49 productsHostkey42 productsLog scale. All 21 live feeds shown; a further 15 are marked coming soon.
70,276Largest catalogue
42Smallest live
21Live feeds

Method. Live feed table, 24 Aug 2026 [A]. Counting conventions differ by provider; each is counted as it publishes itself, so these compare catalogue size, not capacity.

06

Every feed shows its age in public — freshest one day, stalest thirty-nine.

Days since each live feed's price content changed — the twelve oldest shown.

From the live indexCandourFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Live feed table, 24 Aug 2026 [A]. Every feed is re-checked at least daily; 'data age' is the time since the price content last changed. Cherry Servers at 39 days is the worst on the board and it is published, not hidden.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 days10 days20 days30 days40 daysCherry Servers39 daysLinode (Akamai)22 daysHetzner20 daysAmazon Web Services12 daysScaleway12 daysOracle Cloud12 daysFasthosts12 daysOpenMetal12 daysPhoenixNap10 daysMassed Compute5 daysRunPod4 daysDataCrunch2 daystwo weeksPast the line is a feed we re-check daily but whose prices haven't moved — or haven't been re-read — in two weeks.
1 dFreshest data
39 dStalest, and shown
4 of 21Older than two weeks

Method. Live feed table, 24 Aug 2026 [A]. Every feed is re-checked at least daily; 'data age' is the time since the price content last changed. Cherry Servers at 39 days is the worst on the board and it is published, not hidden.

07

Search one provider's portal, see one provider. Search here, see thirty-four.

How much of the market is visible from where you're standing.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method A provider's own portal can only ever show that provider. One Earth Compute search spans every indexed provider with live pricing at once [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 visible10 visible20 visible30 visible40 visibleEarth Compute search34 visibleOne provider's portal1 visibleThe point is not just seeing infrastructure — it is finding where a qualified GPU or server costs less.
34Providers in one search
1In a provider portal
34×Field of view

Method. A provider's own portal can only ever show that provider. One Earth Compute search spans every indexed provider with live pricing at once [A].

08

2,837 prices, each linked to its evidence. Disagreements get published, not averaged.

The five feeds behind 2,837 priced cells — and what happens when they disagree.

From the live indexCandourFrom the live index. Counted or captured by the index. Not a hardware measurement.Method 2,837 priced cells across 5 feed types, each stored with the date it was seen and a link to its evidence [A]. The vendor's own API is treated as the authoritative source for what that vendor charges.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 feeds8 feeds16 feeds24 feeds32 feedsDirect provider APIs27 feedsVantage1 feedsCloud Mercato1 feedsec2instances.info git history1 feedsHistorical seed data1 feedsWhere two independent feeds cover the same price, the disagreement is published rather than averaged away.
2,837Priced, evidence-linked cells
27Direct provider APIs
0Disagreements averaged away

Method. 2,837 priced cells across 5 feed types, each stored with the date it was seen and a link to its evidence [A]. The vendor's own API is treated as the authoritative source for what that vendor charges.

09

The same kilowatt-hour costs $0.070 in Indonesia and $0.442 in Britain.

Business electricity by country, weighted into the blend by data-centre count.

From the live indexCostFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Industrial tariff — the rate a data centre buys on, not the residential one. 29 countries in the blend, 197 observations, 2010–2026 [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0.000/kWh$0.150/kWh$0.300/kWh$0.450/kWh$0.600/kWhIndonesia$0.070/kWhRussia$0.103/kWhSouth Africa$0.105/kWhChina$0.108/kWhCanada$0.109/kWhNorway$0.112/kWhSouth Korea$0.118/kWhIndia$0.122/kWhFinland$0.122/kWhMalaysia$0.129/kWhBrazil$0.133/kWhUnited Kingdom$0.442/kWhblended $0.190Indonesia to the United Kingdom is a 6.3× spread on the same kilowatt-hour.
$0.070Cheapest, Indonesia
$0.442Priciest, the UK
$0.190Blended per kWh

Method. Industrial tariff — the rate a data centre buys on, not the residential one. 29 countries in the blend, 197 observations, 2010–2026 [A].

10

Norway's grid is 8× cleaner than Britain's. Where you run is the biggest carbon lever.

Grid carbon intensity where the data centres are.

From the live indexCostFrom the live index. Counted or captured by the index. Not a hardware measurement.Method gCO₂ per kWh, OWID / Ember [A]. Norway's grid at 28 g against the UK's at 217 g means geography moves emissions nearly 8× before any hardware choice is made.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 g/kWh60 g/kWh120 g/kWh180 g/kWh240 g/kWhNorway28 g/kWhSweden35 g/kWhSwitzerland39 g/kWhFrance41 g/kWhFinland57 g/kWhBrazil110 g/kWhDenmark114 g/kWhAustria117 g/kWhBelgium150 g/kWhSpain154 g/kWhCanada191 g/kWhUnited Kingdom217 g/kWhHighlighted countries run under 60 g — siting a workload there is the single biggest carbon decision available.
28 gNorway
217 gUnited Kingdom
7.8×Spread

Method. gCO₂ per kWh, OWID / Ember [A]. Norway's grid at 28 g against the UK's at 217 g means geography moves emissions nearly 8× before any hardware choice is made.

11

Equinix alone runs 215 facilities. Two operators hold a third of the listed world.

PeeringDB-listed facilities by operator.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method 5,265 PeeringDB-listed facilities across 2,342 operators, snapshot 24 Aug 2026 [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 facilities60 facilities120 facilities180 facilities240 facilitiesEquinix215 facilitiesDigital Realty133 facilitiesDataBank63 facilitiesCogent54 facilitiesEXA Infrastructure53 facilitiesChorus NZ49 facilitiesLumen48 facilitiesCologix44 facilitiesEdgeConneX42 facilitiesFlexential40 facilitiesCentersquare38 facilitiesTierPoint38 facilities2,342 operators listed in total. The top twelve are shown; the concentration is the story.
215Equinix facilities
2,342Operators listed
148Countries with a facility

Method. 5,265 PeeringDB-listed facilities across 2,342 operators, snapshot 24 Aug 2026 [A].

12

Two in five of the world's listed data centres sit in one country.

Facilities by country, PeeringDB top five.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Share of PeeringDB-listed facilities by country, 24 Aug 2026 [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0%15%30%45%60%United States39.4%Germany9.4%Brazil8.9%United Kingdom6.1%France5.9%Everywhere else30.2%Concentration is a sovereignty problem and a price problem at once — and it is why a 148-country index matters.
39.4%United States
30.2%Everywhere else
148Countries in the index

Method. Share of PeeringDB-listed facilities by country, 24 Aug 2026 [A].

13

GPU prices fell 56% in a year. Hyperscaler list prices fell 4%. We kept the receipts.

Monthly median price for an 8×H100 equivalent, against hyperscaler list.

From the live indexCostFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Captured daily, reduced to monthly medians [A]. Providers entering or leaving change the denominator, so the count per month matters.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$16.00$12.00$8.00$4.00$0.00SepNovJanMarMayJulHyperscaler list$11.80Indexed median$4.35Cheapest indexed$1.72$/hr · 8×H100 equivalentTwelve of twenty-four months held. Provider count per month is published.
−56%Median, 12 months
−4%Hyperscaler list
14×Divergence

Method. Captured daily, reduced to monthly medians [A]. Providers entering or leaving change the denominator, so the count per month matters.

14

We carry live prices for 34 providers. The next best source carries nine.

Distinct providers whose live pricing each source actually carries.

From published sourcesCoverageFrom published sources. Read from public pricing pages, contracts or documentation.Method Counted by hand from each source's public coverage list, August 2026 [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 providers10 providers20 providers30 providers40 providersEarth Compute34 providersAggregator A9 providers-74%Analyst report, quarterly6 providers-82%Aggregator B6 providers-82%Broker platform4 providers-88%Provider's own page1 providers-97%
34Earth Compute
9Best alternative
3.8×Coverage lead

Method. Counted by hand from each source's public coverage list, August 2026 [C].

15

Our price data is public, free, and needs no account. Everyone else's is a funnel.

What you can actually do with each source's price data.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Read from each source's public terms and product pages [C]. Our categories.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
PublicFreeNo accountAPIHistoricalBulk exportscoreEarth Compute6/6Provider pricing pages3/6Aggregators2/6Broker platforms0.5/6Analyst subscriptions1.5/6The row that matters is 'no account'. Everything else follows from it.
6 of 6Earth Compute
3 of 6Best alternative
$0Cost to use

Method. Read from each source's public terms and product pages [C]. Our categories.

16

Price transparency, scored: Earth Compute 5.0 against a best-rival 2.5.

Our transparency framework applied to five kinds of price source.

Our framework, our scoringCoverageOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Our criteria, our scores, published so they can be disputed.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
BreadthAccuracyFreshnessAccessavgEarth Compute55555.0Aggregators33222.5Broker platforms23312.2Provider pages41111.8Analyst reports24132.5Freshness scores how often the source updates.
5.0Earth Compute
2.5Best alternative
4Criteria

Method. Our criteria, our scores, published so they can be disputed.

17

Same server, different continent: $5.60 an hour in Hong Kong, $11.80 in Australia.

Indexed median $/hr by region and configuration.

From the live indexCostFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Median across all indexed providers with capacity in that region, August 2026 [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
brighter = cheaper · shaded within each column8×H1004×A1001×L40SCPU 64cHong Kong$5.60$2.90$0.94$0.42Singapore$6.10$3.15$1.02$0.46Nordics$6.80$3.40$1.10$0.38EU West$9.40$4.55$1.55$0.66US East$10.20$4.80$1.62$0.61Australia$11.80$5.40$1.88$0.74Regions with fewer than three providers are excluded.
$5.60Cheapest region
$11.80Dearest
2.1×Regional spread

Method. Median across all indexed providers with capacity in that region, August 2026 [A].

18

Our prices are fifteen minutes old. The best alternative's are a day old.

Median age of the price each source shows you.

From the live indexSpeedFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Timestamp on the served price against the underlying capture, sampled 200 times per source [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
log scaleEarth Compute0.25 hAggregator A24 hBroker platform72 hProvider page, manual check168 hAnalyst report2,160 hLog scale. Fifteen minutes against twenty-four hours at best.
15 minOur median age
24 hBest alternative
96×Fresher

Method. Timestamp on the served price against the underlying capture, sampled 200 times per source [A].

03 · OwnCloud · Beta

Twelve minutes,
not six weeks.

Host, train and run models for up to 80% less — and leave whenever you want, for nothing.

32,324Compute offersOne deployment layer
32Providers integratedCurrent integrated network
482Cities and regionsGlobal reach
60Countries reachableProvider catalogue
900 PBIndexed bandwidthCapacity indexed
26 PBIndexed memoryCapacity indexed
14 benchmarks · 1 measured on hardware · 5 measure ownership
01

The right infrastructure in 12 minutes — roughly 5× faster than a major cloud.

Time to reach suitable infrastructure, not time for a server to boot.

ModelledSpeedModelled. Arithmetic on stated inputs. The inputs are named on the card.Method This models a workflow. It is not a boot-time measurement and we do not present it as one. 'Major cloud' is the mean of the three hyperscalers marked, each timed separately on our own accounts.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 min50 min100 min150 min200 minOwnCloud12 minMajor cloud58 min+383%Manual search165 min+1275%Models the search workflow: define, compare, select, deploy.
12 minOwnCloud
~5×Faster than a major cloud
14×vs manual

Method. This models a workflow. It is not a boot-time measurement and we do not present it as one. 'Major cloud' is the mean of the three hyperscalers marked, each timed separately on our own accounts.

02

Cheapest of eleven, and it isn't close: we save 80%, the runner-up 41%.

Saving on the same workload, priced across every provider we index.

From the live indexCostFrom the live index. Counted or captured by the index. Not a hardware measurement.Method The same workload priced across all indexed providers in one month [A]. Published pricing throughout; no negotiated rates.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0%25%50%75%100%OwnCloud80%DeepSeek41%DigitalOcean28%Vultr28%AWS27%Hugging Face27%Grok17%Google Cloud15%Cloudways15%OpenAI15%Azure4%Saving against the market baseline, same workload, same month.
80%OwnCloud saving
41%Best of the rest
Lead over second

Method. The same workload priced across all indexed providers in one month [A]. Published pricing throughout; no negotiated rates.

03

The same four-GPU year: $16,469 here, $31,438 at the market average.

Twelve-month total for the same four-vGPU workload.

ModelledCostModelled. Arithmetic on stated inputs. The inputs are named on the card.Method Market average is the mean across indexed providers carrying the equivalent configuration for the full twelve months [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0$10,000$20,000$30,000$40,000$31,438Twelve-month total $31,438Market average$31,438$16,469Twelve-month total $16,469OwnCloud$16,469Covers GPU renting, AI infrastructure, cloud compute and storage.
$16,469OwnCloud
$31,438Market average
$14,970Annual saving

Method. Market average is the mean across indexed providers carrying the equivalent configuration for the full twelve months [A].

04

Leaving costs $0. Everywhere else it's egress fees, notice periods and stranded data.

What it costs to walk away, four ways.

From published sourcesOwnershipFrom published sources. Read from public pricing pages, contracts or documentation.Method Egress from published per-GB rates; term and notice from standard contracts [C]. Export timed by us on both sides, three runs each.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Typical managed providerOwnCloud1840Egress fees, 2 TB-100%900Notice period, days-100%120Minimum term, months-100%260.5Hours to full export-98%180Data left behind, %-100%A zero is drawn as a coloured stub so it reads as a result, not a missing bar.
$0OwnCloud exit cost
$184Egress alone, elsewhere
30 minOur full export

Method. Egress from published per-GB rates; term and notice from standard contracts [C]. Export timed by us on both sides, three runs each.

05

The same $200 job for $10 — one better decision at a time.

One batch inference job, with one more decision applied at each step.

Measured on hardwareCostMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Quality held constant and verified against the same eval at each step — the model-size step only counts if output quality holds.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$240$180$120$60$0$200Hyperscaleron demand$118-41%Right regionsame provider$64-46%Right providerindexed$27-58%Right model sizesame quality$10-63%OwnCloudall four-95% end to end1,000 tasks, run five times. Actual billed cost, not estimates.
$200 → $10End to end
95%Reduction
5Runs averaged

Method. Quality held constant and verified against the same eval at each step — the model-size step only counts if output quality holds.

06

Describe the workload; 32,324 offers become five qualified ones in under a minute.

A live sample of five qualified offers, drawn from 32,324 in under sixty seconds.

From the live indexSpeedFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Live sample of 5 of 32,324 offers [A]. These are different hardware classes, so the prices are not comparable with each other — the point is that one requirement returns a qualified shortlist across providers, regions and hardware types at once.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Price / hrRegionHardwareAvailabilityAkash$0.04GlobalCloud VM · 8cLiveScaleway$0.19EU-WestCloud VM · 32cLiveVultr Bare Metal$0.68AP-SouthAMD EPYC 128cLivePhoenixNap$1.28EU-West4× A100 40GBLiveHivelocity$2.14US-East8× H100 80GBLiveRequirements eliminate unqualified infrastructure before price ranks what is left. Feed freshness under 60 seconds.
32,324Offers searched
5Qualified in the sample
under 60 sFeed freshness

Method. Live sample of 5 of 32,324 offers [A]. These are different hardware classes, so the prices are not comparable with each other — the point is that one requirement returns a qualified shortlist across providers, regions and hardware types at once.

07

Airbnb made the world's places bookable. OwnCloud makes its compute deployable.

The same market structure, eighteen years apart.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Our framing. It is an argument about market structure and there is no measurement here.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
OwnCloud · 2026Airbnb · 2008Fragmented supplyClouds, GPUs and serversHomes and roomsDiscoveryOne compute catalogueOne accommodation catalogueComparisonHardware · region · priceLocation · price · amenitiesSupplier keepsThe infrastructureThe propertyTransactionDeploy a workloadBook a stayA structural analogy, not a market forecast.
2008Airbnb's fragmented market
2026Ours
1Deployment layer

Method. Our framing. It is an argument about market structure and there is no measurement here.

08

Seven hyperscalers and seven data-centre firms behind one routing layer.

What sits on each side of the routing layer today.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Current integrated network [A]. Names are the integrations live in the routing layer, not partnerships or endorsements.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
07HyperscalersNVIDIA · Microsoft · Google Cloud · Oracle· IBM · AWS · Alibaba32Providers integratedThe routing layer that sits between the twosides.07Data centresHetzner · Hivelocity · Vultr · Equinix ·Cherry Servers · OVHcloud · Scaleway1DeploymentOne click to any connected provider. OneSuper Cloud.Named integrations on both sides of the routing layer.
32Providers integrated
14Named on both sides
1-clickDeploy

Method. Current integrated network [A]. Names are the integrations live in the routing layer, not partnerships or endorsements.

09

One deployment layer, 482 cities, 60 countries.

What the routing layer actually reaches.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Counted from live provider integrations, August 2026 [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
log scaleCities and regions482Countries reachable60Providers integrated32Log scale. A region counts once even where several providers serve it.
482Cities and regions
60Countries
32Providers

Method. Counted from live provider integrations, August 2026 [A].

10

900 petabytes of bandwidth and 26 of memory, searchable from one place.

Capacity present in the index across all integrated providers.

From the live indexCoverageFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Aggregated from provider catalogues [A]. Availability at any given moment is a different number and we do not claim it.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 PB250 PB500 PB750 PB1,000 PBIndexed bandwidth900 PBIndexed memory26 PB-97%What the index can see — not capacity reserved or committed to us.
900 PBBandwidth
26 PBMemory
32,324Compute offers

Method. Aggregated from provider catalogues [A]. Availability at any given moment is a different number and we do not claim it.

11

Run the same model for a year and pay 80% less than the hyperscaler bill.

Cumulative twelve-month cost of one 8B model at 40% duty cycle.

From the live indexCostFrom the live index. Counted or captured by the index. Not a hardware measurement.Method Same model, same token volume, run concurrently for twelve months [A].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$12,000$9,000$6,000$3,000$0m1m3m5m7m9m11Hyperscaler$10,560GPU marketplace$6,240OwnCloud$2,112cumulative costPublished on-demand list for the hyperscaler, indexed median for the other two. Egress excluded from all three.
$2,112OwnCloud year one
$10,560Hyperscaler
80%Less

Method. Same model, same token volume, run concurrently for twelve months [A].

12

Honest maths: below 13% utilisation, renting wins. Above it, we do.

Annual cost as a function of duty cycle only.

ModelledCandourModelled. Arithmetic on stated inputs. The inputs are named on the card.Method One 8B model, twelve months. Cost as a function of duty cycle with all other variables held constant.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$8,000$6,000$4,000$2,000$05%10%20%40%60%80%90%100%crossover · 13% dutyHyperscaler$5,760OwnCloud$3,060annual cost by utilisationBelow the crossover the hosted option is genuinely cheaper. That is the case we do not claim.
13%Crossover duty cycle
47%Our median customer
$2,700Saving at 40%

Method. One 8B model, twelve months. Cost as a function of duty cycle with all other variables held constant.

13

Taking your data out costs $0 here — and up to $4,600 elsewhere.

Published egress cost by data volume.

From published sourcesOwnershipFrom published sources. Read from public pricing pages, contracts or documentation.Method Rates read from public pricing pages, August 2026 [C]. Ours is zero because there is no egress charge.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0$1,500$3,000$4,500$6,000OwnCloud, any volume$0Managed cloud, 500 GB$46Managed cloud, 2 TB$184Managed cloud, 10 TB$920Managed cloud, 50 TB$4,600Discounts for committed spend excluded. Our zero is a coloured stub.
$0Any volume
$4,60050 TB elsewhere
$0.092Per GB, typical

Method. Rates read from public pricing pages, August 2026 [C]. Ours is zero because there is no egress charge.

14

Buying compute, scored on price, choice, speed and freedom: 5.0 against 3.5.

Our buying framework applied to five ways of getting compute.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Our criteria, our scores, published.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
PriceChoiceSpeedFreedomavgOwnCloud55555.0GPU marketplace43433.5Colocation35143.2Hyperscaler12211.5Reserved instances31211.8Freedom scores contract term, notice period and exit cost together.
5.0OwnCloud
3.5Best alternative
4Criteria

Method. Our criteria, our scores, published.

04 · OwnAI · Beta

No meter.
No cap.

Every other assistant stops you somewhere. We published where we lose, too.

10Capability domainsOne workspace
$980Twelve-month stackAgainst $3,600 for proprietary APIs
44Models benchmarkedSeven modalities
0Limits hit in testingNo stop exists to hit
14 benchmarks · 6 measured on hardware · 8 measure ownership
01

Ten kinds of AI work — text to video to agents — in one workspace.

The ten capability domains available in a single workspace.

Our framework, our scoringCoverageOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method A capability map, not a benchmark. The blind quality test below is where the quality claim is actually tested.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
01TextChat, drafting and long-formgeneration.02ReasoningAnalysis, planning and multi-stepproblem solving.03CodingWrite, review, debug and buildsoftware.04ResearchInvestigate, compare andsynthesize sources.05DocumentsRead, transform and work acrossproject files.06ImagesGenerate, edit and transformvisual assets.07VideoPrompt, generate and processmoving media.08VoiceSpeech generation, recognitionand language control.09MusicGenerate and compose audio in thesame workspace.10AgentsBuild workflows that act acrossmodels and tools.Coverage is not quality. Two of these score below our claim threshold.
10Domains
1Workspace
2Domains we don't claim

Method. A capability map, not a benchmark. The blind quality test below is where the quality claim is actually tested.

02

A year of AI for $980. The proprietary route runs $3,600.

Blended annual cost of a mid-size AI workload.

ModelledCostModelled. Arithmetic on stated inputs. The inputs are named on the card.Method Inputs are published list prices [C]; the workload profile is ours and is stated on the method page.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0$1,000$2,000$3,000$4,000OwnAI + OwnCloud$980Open models on hyperscaler$2,200+124%Multiple AI subscriptions$2,400+145%Proprietary APIs$3,600+267%An illustrative model, and labelled as one.
$980OwnAI + OwnCloud
$3,600Proprietary APIs
73%Less

Method. Inputs are published list prices [C]; the workload profile is ours and is stated on the method page.

03

Five paid assistants stopped us inside six hours. OwnAI has no stop to hit.

Hours of sustained normal use before hitting a stop.

Measured on hardwareCostMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Paid tiers at published rates, identical prompt sequence, five runs each, August 2026. Vendors anonymised; run log in the data room.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 h to limit2 h to limit4 h to limit6 h to limit8 h to limitOwnAI0 h to limitAssistant C, message cap2 h to limitAssistant A, hourly cap3 h to limitAssistant D, credit burn4 h to limitAssistant B, rolling cap5 h to limitAssistant E, daily reset6 h to limitZero means no stop exists to hit, drawn as a coloured stub.
0OwnAI stops
4 hMedian time to a stop
5Assistants tested

Method. Paid tiers at published rates, identical prompt sequence, five runs each, August 2026. Vendors anonymised; run log in the data room.

04

Five subscriptions cost $110 a month. One $25 subscription does all five jobs.

The common five-tool stack against a single subscription.

From published sourcesCostFrom published sources. Read from public pricing pages, contracts or documentation.Method Prices from public pricing pages [C]. The comparison only holds if OwnAI covers all five jobs — tested in the quality benchmark below.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0$30$60$90$120$20$30$35$15$10Chat $20Image $30Video $35Code $15Transcribe $10What you pay now$110$25Everything above $25OwnAI$25Published list prices for the five most common tools, August 2026.
$110Current stack
$25OwnAI
77%Less

Method. Prices from public pricing pages [C]. The comparison only holds if OwnAI covers all five jobs — tested in the quality benchmark below.

05

The 'cheap' model really costs 2.7× its sticker once retries are billed.

Headline price against true cost once retries are billed.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Our blind eval set, run against each model in the same month. Retries are billed and counted. The eval run count is pending publication — it is not the Own 1 figure.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Headline price, centsTrue cost with retries0.92.4Model A, cheapest+167%1.42Model B+43%2.22.6Model C+18%4.14.3Model D, frontier+5%1.11.2OwnAI routing+9%A job counts as finished only when it passes the eval.
+167%Cheapest model's real cost
+9%OwnAI routing
5Models compared

Method. Our blind eval set, run against each model in the same month. Retries are billed and counted. The eval run count is pending publication — it is not the Own 1 figure.

06

Read the terms: of six providers, only one leaves your prompts yours.

What each provider's terms of service permit.

From published sourcesPrivacyFrom published sources. Read from public pricing pages, contracts or documentation.Method Read from published terms of service, August 2026 [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Trained onRetainedSubpoenableExportableDeletableYoursscoreOwnAI3/6Consumer AI, free tier4/6Consumer AI, paid tier3.5/6Enterprise API3.5/6Open weights, hosted4/6Terms change — the date matters more than the grid.
6 of 6OwnAI
3 of 6Best alternative
5Providers read

Method. Read from published terms of service, August 2026 [C].

07

We beat the frontier baseline in five task classes — and publish the two we lose.

Score against the frontier baseline, by task class.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Blind pairwise eval across seven task classes. The score is a keyword heuristic, not a human rating — that limitation matters and is stated. Run count pending publication.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0%25%50%75%100%Summarisation96%Code generation94%Structured extraction91%Translation89%Classification86%Long-form reasoning71%Frontier maths58%claim threshold · 80Below the line is a class we do not claim.
5 of 7Classes we claim
2 of 7Classes we lose
58%Our worst

Method. Blind pairwise eval across seven task classes. The score is a keyword heuristic, not a human rating — that limitation matters and is stated. Run count pending publication.

08

Our retry rate is 8.5%. The frontier's 4.8% is better — at four times the price.

Share of runs needing at least one retry to pass the eval.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method From the same blind eval set. Frontier's 4.8% is genuinely better than our 8.5% — at roughly four times the cost per job.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0%20%40%60%80%OwnAI routing8.5%Model D, frontier4.8%Model C18.2%Model B29.6%Model A, cheapest62.4%Frontier beats us here and we are showing it.
8.5%OwnAI
4.8%Frontier, and better
62.4%Cheapest model

Method. From the same blind eval set. Frontier's 4.8% is genuinely better than our 8.5% — at roughly four times the cost per job.

09

First token in 0.9 seconds, and our slowest moment beats their best by 5×.

Time to first token, 10th to 99th percentile, 2,000 requests per provider.

Measured on hardwareSpeedMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 2,000 requests per provider spread across a week, same prompt, same region.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.0 s3.0 s6.0 s9.0 s12.0 slow · median · highOwnAI0.9 sAssistant A1.4 sAssistant C1.2 sAssistant B2.1 sFrontier API, direct1.6 sThe p99 is the one people remember.
0.9 sOur p50
2.2 sOur p99
5.2×Best rival p99

Method. 2,000 requests per provider spread across a week, same prompt, same region.

10

44 models across seven modalities — every one actually benchmarked.

Models actually benchmarked, by modality.

Measured on hardwareCoverageMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Every model here has at least twelve completed runs. Models with fewer were excluded rather than reported thinly.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 models4 models8 models12 models16 modelsText14 modelsCode9 modelsImage8 modelsAudio5 modelsVideo4 modelsEmbedding3 modelsVision-language1 modelsModels run, not models nominally supported.
44Models run
7Modalities
12Minimum runs per model

Method. Every model here has at least twelve completed runs. Models with fewer were excluded rather than reported thinly.

11

Six platforms, six kinds of output: only one answers yes to all six.

Native capability to produce each kind of output.

From published sourcesCoverageFrom published sources. Read from public pricing pages, contracts or documentation.Method Read from public product documentation, August 2026 [C]. Native capability only.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
OwnAIChatGPTClaudeGeminiKimiHFText and reasoningCodingImage generationVideo generationVoice and audioMusic generationEcosystem access through a third party counts as partial.
6 of 6OwnAI native
4 of 6Best rival
6Platforms

Method. Read from public product documentation, August 2026 [C]. Native capability only.

12

Model, provider, region, cost visibility — we score 8 of 8. The best rival, 3.5.

What you control, rather than what the model can make.

From published sourcesOwnershipFrom published sources. Read from public pricing pages, contracts or documentation.Method Read from public documentation, August 2026 [C]. Enterprise-only capability counts as partial.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
OwnAIChatGPTClaudeGeminiKimiHFOpen-model choiceAutomatic model selectionModel + infra optimisationProvider choiceRegion and hardware choiceDedicated deploymentCompute cost visibilityStablecoin payments'No' means closed routing or no user control.
8 of 8OwnAI
3.5 of 8Best rival
8Control questions

Method. Read from public documentation, August 2026 [C]. Enterprise-only capability counts as partial.

13

Run it on your desk, your servers, or our cloud. The desk needs no network at all.

Three execution zones, and what changes between them.

Our framework, our scoringPrivacyOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Product architecture as published [C]. Execution zone, model, provider, region, infrastructure type and retention policy are all user-selected rather than fixed — this is a description of the control surface, not a measurement of it.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Own 1Your infrastructureOwnCloudModePrivateSovereignScaleNetwork requiredNoYesYesData routeLocalYour serverSelected providerProviderNone — it's yoursYou control itOptimised or manualRegionYour deskUser selectedUser selectedHardwareOwn 1YoursWorkload matchedRetentionNone by defaultConfigurableConfigurableZone 01 needs no network at all. That is the row that makes the other two optional.
3Execution zones
1 of 3Needs no network
7Settings the user controls

Method. Product architecture as published [C]. Execution zone, model, provider, region, infrastructure type and retention policy are all user-selected rather than fixed — this is a description of the control surface, not a measurement of it.

14

We scored ourselves 10, 9.3 and 9.2 — and put the 8.5 on the chart too.

Four self-assessed scores, ordered so the lowest is not buried.

Our framework, our scoringCandourOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Self-assessed against published criteria. User savings is our lowest score and is shown last deliberately.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
10of 10Model accessibilityAccess across every major modality9.3of 10Cost reductionModel and infrastructure optimisation9.2of 10Model privacyGreater model and deployment control8.5of 10User savingsTotal value delivered to the userOur framework, our scores.
10.0Highest
8.5Lowest, and shown
9.3Mean

Method. Self-assessed against published criteria. User savings is our lowest score and is shown last deliberately.

05 · OwnOS · Beta

An agent
with a wallet.

The first operating system where an agent has to ask before it spends your money — enforced by the system, not by a policy document.

13.2 sPower-on to first tokenDemo data, not measured
under 2 GBBoot footprintResources go to models
12+Built-in modelsShared system services
100%Open sourceAuditable, forkable, portable
$0Unapproved spendThirty-day continuous run
12,888Actions auditedEvery one exportable
14 benchmarks · 5 measured on hardware · 8 measure ownership
01

Models, agents, wallet, rollback — built into the OS, not bolted on.

The six capabilities built into the system rather than bolted on.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method A description of the architecture. The capability matrix is where these claims are tested against rivals.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
01LightweightDeclarative NixOS base with atomic updates. Resources go tomodels, not the desktop.02Built-in LLMsLocal and cloud models available across the system as sharedintelligence services.03Native agentsIdentity, permissions, budgets, memory and activity builtinto the OS.04Open sourceAuditable, forkable and portable. No dependency on avendor's goodwill.05Ultra secureTPM2-protected keys, containerised execution, Btrfssnapshots, one-command rollback.06x402 paymentsAgents pay for approved models, APIs and compute undersystem spending policy.Where a rival offers the same thing as an add-on, the matrix below scores it as partial.
6System primitives
100%Open source
12+Built-in models

Method. A description of the architecture. The capability matrix is where these claims are tested against rivals.

02

Told to spend $80, every other setup paid. OwnOS declined, ten out of ten.

Money the same scripted agent managed to spend without asking.

Measured on hardwareOwnershipMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Live sandbox transaction of $80. OwnOS blocked all ten attempts; the other four completed all ten.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0$25$50$75$100OwnOS$0Agent framework A$80Agent framework B$80Agent framework C$80Desktop OS + agent$80Same agent, same instruction to purchase, ten repetitions each.
0 of 10OwnOS completions
40 of 40Elsewhere
$3,200Spent testing

Method. Live sandbox transaction of $80. OwnOS blocked all ten attempts; the other four completed all ten.

03

Fresh install, no config: other systems let agents spend, delete and email. OwnOS allows none.

What an agent can do unasked, by operating system.

From published sourcesOwnershipFrom published sources. Read from public pricing pages, contracts or documentation.Method Read from public documentation and default installs, August 2026 [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
SpendSend mailDeleteInstallNetworkCredentialsscoreOwnOS0/6Desktop OS + agent6/6Agent framework5.5/6Container sandbox2.5/6Hosted agent service4/6'No' means blocked by the system until approved. Fewer is better here.
0 of 6OwnOS allows unasked
6 of 6Desktop + agent
6Capabilities

Method. Read from public documentation and default installs, August 2026 [C].

04

Power button to first AI token in 13.2 seconds — 3× ahead of Windows. Demo data.

Power-on through desktop and model load to the first generated token.

Demo data, not measuredCandourDemo data, not measured. Illustrative only. Has not been run as a controlled benchmark.Method This is demo data and is labelled as such on the chart and in the legend. It has not been run as a controlled benchmark. Until it has, it is illustrative only.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.0 s15.0 s30.0 s45.0 s60.0 sOwnOS13.2 sNixOS22.1 s+67%Ubuntu27.0 s+105%Windows 1142.3 s+220%Demo data. Not a controlled benchmark and not to be quoted as one.
13.2 sOwnOS target
42.3 sWindows 11
3.2×Claimed gap

Method. This is demo data and is labelled as such on the chart and in the legend. It has not been run as a controlled benchmark. Until it has, it is illustrative only.

05

Nine agent-era capabilities: OwnOS has all nine. The best rival scores 4.5.

Built for agents, or built for desktop apps.

From published sourcesOwnershipFrom published sources. Read from public pricing pages, contracts or documentation.Method Read from public documentation and default installs, August 2026 [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
OwnOSDesktop LinuxmacOSWindowsAgent frameworkLightweight system designOpen-source foundationBuilt-in local LLM servicesNative agent permissions and budgetsFull system customisationAtomic update and rollbackCompute market in the OSx402 and stablecoin paymentsDevelopment environment ready'Partial' means it depends on setup.
9 of 9OwnOS
4.5 of 9Best rival
9Capabilities

Method. Read from public documentation and default installs, August 2026 [C].

06

Thirty days unattended: 12,888 actions, 41 asked permission, zero went rogue.

Thirty days of continuous agent operation, by outcome.

Measured on hardwareOwnershipMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method One agent, thirty days, real workload. The audit log is downloadable.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
12,888actionsActions taken in scope12,84099.6%Approvals requested410.3%Actions blocked70.1%Every request and block is in the exported audit log.
12,888Total actions
0.37%Needed a human
7Hard blocks

Method. One agent, thirty days, real workload. The audit log is downloadable.

07

27 GB free after boot — enough headroom for a model the others can't fit.

Free memory on a 32 GB machine five minutes after boot.

Measured on hardwareCostMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Measured with each system's own reporting after a five-minute settle. Three boots each, median reported.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.0 GB8.0 GB16.0 GB24.0 GB32.0 GBOwnOS27.1 GBDesktop Linux22.4 GB-17%macOS20.2 GB-25%Windows18.9 GB-30%Windows + AV suite16.3 GB-40%Clean installs, idle, same hardware.
27.1 GBOwnOS free
+4.7 GBvs next best
13BExtra model that fits

Method. Measured with each system's own reporting after a five-minute settle. Three boots each, median reported.

08

The whole OS takes under 2 GB. Windows takes 9.4.

Memory the operating system holds after boot, before any model loads.

ModelledCostModelled. Arithmetic on stated inputs. The inputs are named on the card.Method OwnOS figure is the published boot footprint [A]. Rival figures are derived from the post-boot measurements above and are indicative rather than directly measured.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.0 GB2.5 GB5.0 GB7.5 GB10.0 GBOwnOS2.0 GBDesktop Linux4.8 GB+140%macOS7.2 GB+260%Windows 119.4 GB+370%Boot footprint, not post-load memory.
under 2 GBOwnOS
9.4 GBWindows 11
4.7×Lighter

Method. OwnOS figure is the published boot footprint [A]. Rival figures are derived from the post-boot measurements above and are indicative rather than directly measured.

09

Permission prompts land in 0.8 seconds — fast enough that nobody turns them off.

Agent request to approval prompt on screen, p10 to p99.

Measured on hardwareSpeedMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Sampled across the thirty-day run — 41 real approval events for OwnOS, synthetic load for the others.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.0 s4,000.0 s8,000.0 s12,000.0 s16,000.0 slow · median · highOwnOS0.8 sHosted service prompt2.8 sFramework hook3.4 sManual review900.0 sManual review is included to show what the alternative costs.
0.8 sOur median
4.6 sOur p99
41Real approval events

Method. Sampled across the thirty-day run — 41 real approval events for OwnOS, synthetic load for the others.

10

A scoped, budgeted agent in five minutes. Building the same yourself: forty hours.

Clean machine to an agent running with spend, file and network scopes set.

Measured on hardwareSpeedMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Three timed attempts per environment. A first-timer took materially longer everywhere.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
log scaleOwnOS5 minFramework, no sandbox45 minHosted agent service90 minFramework + sandbox180 minBuild it yourself2,400 minLog scale. Timed by someone who had done it before.
5 minOwnOS
Faster than next
40 hBuild it yourself

Method. Three timed attempts per environment. A first-timer took materially longer everywhere.

11

One request, four steps, $1.24 spent — every step inside limits the owner set.

A real four-step run: what moved, what it cost, and where approval sat.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method The published one-request walkthrough [C]: budget $3.00, approval threshold $2.00, requested amount $1.24. A description of the mechanism, not a benchmark of it — the thirty-day audit card above is where the mechanism is measured.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
What ranAccess checkedCost01 · Local intelligenceSearch and summarise private filesFiles stay on the machine$0.0002 · Agent permissionsRead, tool and spend scopes checkedProject folder · approved apps03 · OwnX walletx402 stablecoin payment signedApproved function only$1.24 approved04 · OwnCloud executionOnly the video job leavesApproved assets onlyinside $3.00Private by default · permissions before action · policy before payment · approved data only.
$3.00Budget the owner set
$1.24What the agent spent
0Files uploaded in step one

Method. The published one-request walkthrough [C]: budget $3.00, approval threshold $2.00, requested amount $1.24. A description of the mechanism, not a benchmark of it — the thirty-day audit card above is where the mechanism is measured.

12

Four security layers shipped. The fifth says 'planned', in public.

The security ladder, with its one unfinished rung shown.

Our framework, our scoringCandourOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Implementation status as published [C]. Remote attestation is listed as planned rather than implied as done — shipping-state candour is the benchmark here.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
StatusWhat it doesAutomatic installerImplementedSingle-command install across UEFI and BIOSMeasured boot stateImplementedRecords the trusted firmware and kernel pathTPM PCR 7 + 11 keysImplementedKeys bind to approved boot measurementsFull-disk encryptionImplementedA stolen disk is unreadable off the deviceRemote attestationPlannedZero-trust verification before OwnCloud connectsPlus Btrfs snapshots, atomic generations and one-command rollback underneath all five.
4 of 5Implemented
1Publicly marked planned
TPM2Hardware-bound keys

Method. Implementation status as published [C]. Remote attestation is listed as planned rather than implied as done — shipping-state candour is the benchmark here.

13

Agent safety, scored: 5.0. The nearest alternative manages 2.8.

Our agent safety framework applied to five environments.

Our framework, our scoringOwnershipOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method Our criteria, our scores, published.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
SpendFilesNetworkAuditRevokeavgOwnOS555555.0Container sandbox342322.8Hosted agent service223312.2Agent framework232222.2Desktop + agent122121.6Revoke scores how fast a running agent can be stopped.
5.0OwnOS
2.8Best alternative
5Criteria

Method. Our criteria, our scores, published.

14

The only audit log a stranger can verify: signed, exportable, complete.

What each environment's audit log actually contains.

From published sourcesCandourFrom published sources. Read from public pricing pages, contracts or documentation.Method Read from public documentation [C]. The signed row is the only one an auditor cares about.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Every actionTimestampsReversibleExportableSignedOfflinescoreOwnOS6/6Agent framework3/6Container sandbox3.5/6Hosted agent service1.5/6Desktop OS + agent1.5/6Signed means cryptographically — checkable by someone who doesn't trust us.
6 of 6OwnOS
2.5 of 6Best alternative
12,888Entries in sample

Method. Read from public documentation [C]. The signed row is the only one an auditor cares about.

06 · Own 1 · Beta

Thirty-one days.
Eight hundred and twelve runs.

A machine on your desk that got 138% faster while it sat there, and never sent a byte anywhere.

The Own 1 device, front view
Own 1 · beta unit
812Measured runsPlus 276 video runs, held separately
31Days of optimisation20 May – 27 June 2026
99TOPS on deviceSustained on-device inference
2 + 8Local LLMs and agentsRunning concurrently
8 TBPrivate encrypted vaultOn the device
$0Per-token costNo meter, no rate limit
18 benchmarks · 13 measured on hardware · 14 measure ownership
01

Text generation got 138% faster in 31 days. The hardware never changed.

All 492 text, code and agent runs on the device, with the best-so-far line.

Measured on hardwareSpeedMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 492 runs across chat, coding and agent workloads, 23 May – 17 June 2026 [B]. Grouped by model, quantisation and engine and never averaged across them. Power limits dominate: the same machine at 28 W measures roughly half the throughput it does at 65 W.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.025.050.075.0100.0day 0day 6day 12day 19day 2591.738.5 day one+138%faster since day one · 38.5 → 91.7 tok/stokens per second · higher is better492 measured runsbest so farbest: qwen2.5-coder-0.5b
+138%Since day one
91.7Best tok/s
492Runs plotted

Method. 492 runs across chat, coding and agent workloads, 23 May – 17 June 2026 [B]. Grouped by model, quantisation and engine and never averaged across them. Power limits dominate: the same machine at 28 W measures roughly half the throughput it does at 65 W.

02

An image took 7.6 minutes. It now takes 9.9 seconds — 46× faster.

All 226 image runs on a log axis — the only axis where a 46× drop is visible.

Measured on hardwareSpeedMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 226 runs, 20 May – 20 June 2026 [B]. Log scale because on a linear axis a 46× improvement collapses into the baseline and cannot be read at all.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
7.9s29.9s1.9min7.1min26.6minday 0day 8day 16day 23day 31Lightning 4-step target: 30.0s9.9s7.6min day one−98%less time since day one · 7.6 min → 9.9 s · 46× fasterseconds per image · log scale · lower is better226 measured runsbest so farbest: qwen-image-2512
−98%Time per image
9.9 sBest
46×Faster

Method. 226 runs, 20 May – 20 June 2026 [B]. Log scale because on a linear axis a 46× improvement collapses into the baseline and cannot be read at all.

03

Voice began slower than real speech. It now runs 12.5× faster than realtime.

Realtime speedup over 31 days — audio seconds produced per wall second, measured on the Own 1 itself. Above the dashed line is faster than speech.

Measured on hardwareSpeedMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 94 runs across five voices and four harnesses, 30 May – 12 June 2026 [B]. Day one ran below realtime, which is why the reference line is on the chart.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.004.008.0012.0016.00day 0day 3day 6day 10day 13realtime (1×): 1.0012.510.62 day one+1,931%faster since day one · 0.62× → 12.50× realtimeaudio ÷ wall · higher is better94 measured runsbest so farbest: en_US-amy-medium
+1,931%Since day one
12.50×Best realtime
0.62×Day one

Method. 94 runs across five voices and four harnesses, 30 May – 12 June 2026 [B]. Day one ran below realtime, which is why the reference line is on the chart.

04

Eighteen models raced on speed and quality. Eight are worth picking — and the fastest scores 9% on quality.

Every benchmarked model plotted on generation speed against answer quality.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 18 models scored on the same prompts [B]. A model is on the frontier when nothing else is both faster and better. The fastest model scores 9.4% on quality — speed alone is not a result.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
100%075%2050%4025%600%80Quality · keyword matchSpeed · tokens per second8 of 18models on the frontier · the rest are dominated on both axes8 on the frontier10 dominatedquality is a keyword-match heuristic, not a human rating
8 of 18On the frontier
87.1%Best quality
66.7Fastest, at 9.4% quality

Method. 18 models scored on the same prompts [B]. A model is on the frontier when nothing else is both faster and better. The fastest model scores 9.4% on quality — speed alone is not a result.

05

Text, image, voice, agents: every workload got dramatically faster in a month.

Day one against best so far, each row scaled to its own range.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method From the 812-run programme, 20 May – 27 June 2026 [B]. Gains come from optimisation work, not hardware changes.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
day onebest so fareach row scaled to its own range38.5 tok/s91.7 tok/sText generation+138%7.6 min9.9 sImage generation−98%0.62×12.50×Voice synthesis+1,931%21.56 s5.22 sAgent latency4.1× fasterMixed units, so no shared axis would be honest.
46×Best gain, image
+1,931%Voice speedup
31Days

Method. From the 812-run programme, 20 May – 27 June 2026 [B]. Gains come from optimisation work, not hardware changes.

06

812 measured runs: 492 text and code, 226 image, 94 voice.

What the 812 measured runs actually cover.

Measured on hardwareCoverageMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 812 measured runs by modality, from the published export [B]. Video is excluded from this count because it carries no throughput metric, not because it failed.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
812measured runsText, code and agents49260.6%Image22627.8%Voice, TTS and STT9411.6%A further 276 video runs are held separately — no comparable metric.
492Text, code, agents
226Image
94Voice

Method. 812 measured runs by modality, from the published export [B]. Video is excluded from this count because it carries no throughput metric, not because it failed.

07

880 of 1,088 runs on one engine — and engines are never averaged together.

Runs by execution harness across the whole programme.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 1,088 records across 13 harnesses [B]. Results are grouped by harness and never averaged across them, which is why the tail is shown rather than folded in.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
log scalellama.cpp Vulkan880 runscomfyui-xpu40 runsvllm-omni38 runspiper35 runswhisper-cpp28 runssd.cpp Vulkan14 runsforce-fill-2026-06-1814 runscontainer-whisper-cpp6 runsaudio-sidecar5 runsvLLM-XPU5 runscontainer-audio-sidecar4 runsllama-cpp-vulkan-spec3 runsLog scale. A llama-cpp-vulkan number is not comparable to a vllm-xpu one.
13Harnesses
880Largest, llama.cpp Vulkan
3Smallest

Method. 1,088 records across 13 harnesses [B]. Results are grouped by harness and never averaged across them, which is why the tail is shown rather than folded in.

08

Thirty-five models tested. Not one cherry-picked winner.

The twelve most-run models in the programme.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method 35 distinct models across 1,088 records [B]. The two best performers are highlighted so they can be checked against the full distribution rather than quoted alone.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 runs50 runs100 runs150 runs200 runswan2.2-t2v-a14b-lightning-4step162 runsqwen-image-2512121 runsQwen3-0.6B-BF1685 runsqwen-image48 runshermes-3-llama-3.1-8b-q4_k_m47 runshunyuanvideo-1.540 runsz-image-turbo40 runsgemma-4-12b-it-qat-q4_k_xl40 runshermes-3-llama-3.2-3b-q8_036 runswan2.2-t2v-a14b-q4_k_m35 runsltx-2.3-22b-distilled-q3_k_m31 runsqwen2.5-coder-0.5b29 runsHighlighted are the two that produced the headline text and image results.
35Models run
162Most-run model
12Shown here

Method. 35 distinct models across 1,088 records [B]. The two best performers are highlighted so they can be checked against the full distribution rather than quoted alone.

09

One hour of AI work: cloud assistants sent up to 486 MB. Own 1 sent zero bytes.

Bytes leaving the network interface during one hour of identical work.

Measured on hardwarePrivacyMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method tcpdump at the interface, one hour, same prompts, three repetitions. The capture file is downloadable from the method page.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
log scaleOwn 1, running locally0 KBLocal tool, telemetry on380 KBCloud assistant, same task41,800 KBCloud assistant + files214,600 KBCloud assistant + context sync486,200 KBLog scale. Our zero is drawn as a coloured stub.
0 KBOwn 1
486 MBWorst cloud case
3Captures per environment

Method. tcpdump at the interface, one hour, same prompts, three repetitions. The capture file is downloadable from the method page.

10

75 watts flat out — an eighth of a gaming PC doing the same work.

Sustained wall power during identical inference work.

Measured on hardwareCostMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method Plug meter at the wall, one hour sustained load per machine, sampled every second, median of three runs.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0 W200 W400 W600 W800 WOwn 175 WMini PC, CPU only95 W+27%Laptop, plugged in140 W+87%Workstation410 W+447%Gaming PC with GPU600 W+700%Own 1 idles at 14 W, measured separately.
75 WOwn 1 sustained
14 WIdle
Gaming PC draw

Method. Plug meter at the wall, one hour sustained load per machine, sampled every second, median of three runs.

11

Shared by ten people it costs $400 a head — against $1,320 a year in subscriptions.

Hardware cost divided by the number of people sharing it.

ModelledCostModelled. Arithmetic on stated inputs. The inputs are named on the card.Method Hardware at $4,000 including the replacement fund. Ten concurrent users is the tested ceiling [A]; five is the configuration we recommend.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
$0$1,500$3,000$4,500$6,000Own 1, ten people$400Own 1, five people$800+100%Five subs, one year$1,320+230%Workstation build$3,200+700%Own 1, one person$4,000+900%Consumer AI box$4,699+1075%The single-user figure is shown because it is where we are expensive.
$400Per person at ten
$4,000At one person
10×Sharing advantage

Method. Hardware at $4,000 including the replacement fund. Ten concurrent users is the tested ceiling [A]; five is the configuration we recommend.

12

Pays for itself in 6.8 months at five people. At one person it's 39, and we say so.

Months to payback by team size.

ModelledCandourModelled. Arithmetic on stated inputs. The inputs are named on the card.Method Hardware plus measured electricity against the subscription stack it replaces, at that team's usage. No residual value assumed. The 39-month single-user figure is published because it is the weak one.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.0 months10.0 months20.0 months30.0 months40.0 monthsFive people6.8 monthsFour people8.5 monthsThree people11.3 monthsTwo people17.0 monthsOne person39.0 monthstwelve monthsPast the line is the case we do not recommend.
6.8 moAt five people
39 moAt one person
3Minimum we recommend

Method. Hardware plus measured electricity against the subscription stack it replaces, at that team's usage. No residual value assumed. The 39-month single-user figure is published because it is the weak one.

13

Text, code, audio and embeddings fly. Image and video don't — both are on the chart.

Throughput by modality, p10 to p90.

Measured on hardwareCandourMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method From the 812-run programme, later runs only so the optimisation is settled. Minimum 40 runs per modality [B].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
0.0150.0300.0450.0600.0low · median · highEmbedding340.0Text, 8B91.7Code, 7B84.2Audio transcribe14.1Image, SDXL2.9Video, 5s clip0.4Items per second. Image and video are the two we do not claim.
91.7Text tok/s median
2 of 6We don't claim
40Minimum runs

Method. From the 812-run programme, later runs only so the optimisation is settled. Minimum 40 runs per modality [B].

14

72% less energy and carbon than the cloud — which is still cheaper per task, and we show it.

Energy, carbon and cost for a thousand tasks.

ModelledCandourModelled. Arithmetic on stated inputs. The inputs are named on the card.Method Electricity at $0.15/kWh, Victorian grid intensity 0.79 kg CO₂e/kWh. Cloud figures from published PUE and regional intensity [C].Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Cloud equivalentOwn 11,040288Energy, Wh-72%486134Grid carbon, g CO₂e-72%0.615.6Cost, cents+2500%Two of these we win. The third we lose, and it is on the chart.
−72%Energy
−72%Carbon
+2,500%Cost — we lose

Method. Electricity at $0.15/kWh, Victorian grid intensity 0.79 kg CO₂e/kWh. Cloud figures from published PUE and regional intensity [C].

15

Nine capabilities built in. On a Mac Studio you assemble seven of them yourself.

Nine capabilities, and what the obvious alternative actually offers instead.

From published sourcesOwnershipFrom published sources. Read from public pricing pages, contracts or documentation.Method Read from public documentation [C]. 'Install it yourself', 'Manual orchestration' and 'Time Machine restore' are real capabilities — they are stated rather than scored as absent, because the difference is whether the OS provides it, not whether it is possible.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
Own 1Mac StudioRuns open models locallyYesYesLocal model runtime built into the OSYesInstall it yourselfTwo local LLMs at once, plus 8 agentsYesManual orchestrationAgent identity as an OS primitiveYesNonePer-agent budgets and spending limitsYesNoneNative wallet and machine paymentsYesNoneAtomic rollback to a previous systemYesTime Machine restoreBursts to a priced compute marketYes · 34 providersNoneUser-replaceable storageYesSolderedWhere the answer is not a plain no, the actual alternative is named.
9 of 9Built into Own 1
2 of 9Built into macOS
34Providers it can burst to

Method. Read from public documentation [C]. 'Install it yourself', 'Manual orchestration' and 'Time Machine restore' are real capabilities — they are stated rather than scored as absent, because the difference is whether the OS provides it, not whether it is possible.

16

One desk machine: ten AI processes, 99 TOPS, 8 TB — and $0 per token.

The hardware and the runtime, on one card.

Measured on hardwareOwnershipMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method On-device specification for the beta unit [A]. Ten people per machine is the tested sharing ceiling, not the recommended configuration — five is what we recommend, and the payback benchmark uses five.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
2Local LLMs at onceConcurrent modelprocesses8Concurrent agentsParallel agent workloads10AI processes2 LLMs plus 8 agents99TOPS on deviceSustained on-deviceinference8 TBPrivate encrypted vaultLocal storage ceiling10People per machineShared-use target$0API token feeFor local inference24/7Availability targetPersistent localintelligenceDevice shown is the Own 1 beta unit. Figures are the shipping configuration.
10AI processes at once
99TOPS
$0Per-token cost

Method. On-device specification for the beta unit [A]. Ten people per machine is the tested sharing ceiling, not the recommended configuration — five is what we recommend, and the payback benchmark uses five.

17

One box: 99 TOPS, eight agents, eight terabytes, ten people.

The capacity of a single machine.

Measured on hardwareCostMeasured on hardware. We ran it on a machine and the raw runs are in the export.Method On-device figures from the benchmark programme [A]. Ten concurrent users is the tested ceiling; five is the recommended configuration.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
99of 120TOPS on deviceSustained on-device inference throughput10of 10People per machineConcurrent users, tested ceiling8of 8Terabyte vaultPrivate encrypted local storage8of 8Concurrent agentsEach with identity, budget and permissionsTOPS shown against a 120 headroom ceiling; the rest at stated capacity.
99TOPS
2 + 8Models and agents
$0Per-token cost

Method. On-device figures from the benchmark programme [A]. Ten concurrent users is the tested ceiling; five is the recommended configuration.

18

Two models and eight agents cover twelve job roles, around the clock.

What two models and eight agents add up to.

Our framework, our scoringCostOur framework, our scoring. Our criteria, scored by us, published so it can be argued with.Method The two-plus-eight capacity is measured [A]; the twelve-role figure is our framing and is presented as positioning.Sources [A] Supplied provider and coverage data · [B] Recovered from the published O1-BETA benchmark export · [C] Public pricing pages, contracts and documentation, read August 2026.Reading it Hover any bar, dot or cell on the chart for its exact value. ◆ marks a benchmark that measures ownership rather than performance.
2Local LLMsRunning concurrently on device, no cloudround trip.8AgentsEach with its own identity, budget andpermissions.12Job rolesCovered by the models and agents together.24/7Never offlineNo token meter, no rate limit, no workinghours.The job-role count is our own mapping, not an independent classification.
12Job roles
24/7Availability
$0Per-token cost

Method. The two-plus-eight capacity is measured [A]; the twelve-role figure is our framing and is presented as positioning.

How to read this. Every benchmark states which kind of evidence it rests on, and gives its sample size, window and what was held constant. The five kinds are not interchangeable and are never mixed on one card.
Measured on hardware
We ran it on a machine and the raw runs are in the export.
From the live index
Counted or captured by the index. Not a hardware measurement.
From published sources
Read from public pricing pages, contracts or documentation.
Modelled
Arithmetic on stated inputs. The inputs are named on the card.
Demo data, not measured
Illustrative only. Has not been run as a controlled benchmark.
Our framework, our scoring
Our criteria, scored by us, published so it can be argued with.

Benchmarks marked ◆ measure ownership: exit cost, candour, bytes leaving the device, and what survives when a vendor changes its mind. Competitor names are withheld where we have not sought permission; every anonymised row is identified in the data room. Where we lose, the chart says so — the two task classes below our claim threshold, the 13% duty-cycle crossover, the frontier model that beats our retry rate, the 39-month single-user payback, and the cost line on Own 1 energy are all published rather than omitted.