The Decentralized
Man

LOG-008 ·

Forecast: By 2030, the Best Model You Use Daily Will Run in Your House

A falsifiable prediction about local AI, with the three numbers that will decide it and the confidence I am prepared to be graded on.

Words
701
Est. read
3.0 min
Confidence
0.72
Topics
forecast, compute, local-first

I make forecasts with numbers attached so that future me can grade present me. Today's is about where intelligence will physically live. Stated in full, so there is no wriggle room later:

By 2030-12-31, the language model I personally invoke most often on a typical day will execute on hardware inside my house. Confidence: 0.72.

Note what this does not claim. Not that frontier models will run locally (they will not; the largest training runs will stay industrial for reasons of physics and capital). Not that cloud AI shrinks in absolute terms (it will grow). The claim is about the workhorse: the model that answers the routine question, drafts the routine paragraph, triages the routine inbox. Median invocation, not maximum capability.

The three numbers that decide it

Number one: the capability lag. How many months behind the frontier is the best model that runs on a 32 GB consumer machine? My logs of published benchmarks put that lag at roughly 30 months in 2023 and roughly 12 months today. The frontier keeps moving, but the distillation pipeline moves faster. If the lag drops below about 6 months, the cloud's quality argument stops applying to routine work, because routine work does not need this quarter's frontier. It needs last year's, which will be sitting in my office drawing 45 W.

Number two: memory per dollar. Local inference is bottlenecked less by raw compute than by fast memory. Consumer machines with unified memory in the 128 GB class have fallen in price by roughly half in three years. Extrapolation is a dangerous habit, so I will only say: the trend does not need to continue heroically. It needs to continue moderately.

Number three: the privacy premium going negative. Today, local costs more than cloud per token of comparable quality. But a cloud query's true price includes the retention policy you did not read. The moment local-equivalent quality costs the same as cloud at the point of use, the privacy tiebreaker decides it, and defaults follow tiebreakers. Watch what the operating system vendors preinstall; the OS default is where this contest actually ends.

The base rate argument

This is the third time computing has made this exact round trip. Mainframe terminals became personal computers when the economics crossed. Client-server became the cloud when bandwidth beat local administration. Every reversal happened when the decentralized option became merely adequate, not superior. Adequacy plus autonomy beats excellence plus dependency for the median task; the entire history of the PC is that sentence.

Notice also who is funding the local wave: the same device manufacturers who lost the last decade to the cloud. Silicon roadmaps are public commitments, and every major one now allocates die area to neural inference. Chip fabs are a three-year lookahead. The bet is already placed by people with better information than me.

What would falsify me

I keep a written list, reviewed quarterly:

Signal Bearish threshold
Capability lag (frontier vs 32 GB local) Widens past 18 months for 4 straight quarters
Frontier training cost curve Keeps steepening, and quality scales with it for routine tasks
OS-default assistants Still cloud-routed for >80% of queries in 2029
My own usage log Cloud share of my invocations rising year-over-year

That last row matters most. I log which model answers my queries the way I log my espresso ratios, and today the local share of my personal invocations is 41%, up from 9% eighteen months ago. One man's logs are an anecdote, but they are at least an honestly kept anecdote.

Why 0.72 and not higher

Because two forces push the other way and I decline to pretend otherwise. First, agents that act on your behalf may need to live near the data they act on, and your email regrettably lives in the cloud. Second, subsidized cloud inference could stay below cost for years; venture capital has outlasted better theses than mine.

So: 0.72. Not a conviction, a position, sized accordingly. On 2031-01-01 I will publish the grade next to this entry. If I am wrong, the entry stays up. Deleting your misses is the one form of centralization I could actually achieve, and declining to do it is the point of keeping the log in public.