Mastermind · July 15, 2026
Buckets and Taps

In May I wrote Buckets for Rainwater. The argument was that the next decade of software value starts with catchment: place a bucket where the data is freshest, where nobody else can stand, and let the rain do the rest. I still believe every word of it.
But it was half a thesis. A bucket full of rainwater is not a business. Nobody has ever paid a water bill for a reservoir. They pay for what comes out of the tap.
So this is the other half, and the full frame I now use for where AI value gets captured. I call it Buckets and Taps. BAT, if you want the shorthand. It is part investment memo, part personal manifesto, because I am both betting on this frame and building inside it.
The bucket, briefly
The bucket is the capture layer. Right now the most valuable raw material in the world — human intelligence in motion — is evaporating. Decisions made in meetings, prompts sent to agents, voice memos, drafts, the half-formed idea in a WhatsApp note at 2am. It is scattered across twenty tools and forty data formats, disparate and fragmented, owned by nobody because it lands nowhere.
The bucket's job is centralization: pull those fragmented streams into one data plane, under one unified ontology, so that a fact captured in a meeting and a fact captured in a codebase can actually be joined. Placement decides everything. A bucket under the public web catches the same storm-drain runoff everyone else has. A bucket upstream of a private workflow catches snowmelt only you can touch.
That was the May essay. Here is what it missed.
Full buckets that never paid anyone
The graveyard of the last data era is full of buckets. Enterprises spent a decade and hundreds of billions building data lakes, and most of them became data swamps: petabytes in, nothing out. The data was collected. It was never drinkable.
The investor consensus has caught up to this. The 2026 version of the data-moat debate is notably more conditional than the 2017 version: static proprietary data is no longer treated as a moat by itself. What investors now underwrite is the active flywheel — a closed loop where usage generates data, data improves the product, and the improved product generates more usage. A bucket with no outflow has no loop. It is not a moat. It is a lake.
This is the correction to my own thesis. The bucket is necessary. It is not sufficient. Water that sits, stagnates.
The tap
The tap is everything that turns stored water into pressure at the point of use. You have the bucket collecting all the rainwater. Now you need the system that squeezes it, massages it, and pours it out as something a human — or increasingly, another agent — can actually drink. The bucket collects rainwater. The tap makes it flow out as gold.
When I look at what is actually working in 2026, the tap has three parts, stacked.
The first part is meaning. Before data can flow, it has to mean the same thing everywhere. This is the unified ontology, and the clearest proof it matters is Palantir. Their entire architecture bets that the valuable object is not the data but the decision, and that an ontology which unifies fragmented ERPs, CRMs, sensors, and documents into shared semantic objects is what lets AI actually touch operations. The 2026 enterprise-agent literature keeps landing on the same diagnosis: agents fail in production because meaning fragments across systems and business definitions live as tribal knowledge — a context problem, not a model problem. The ontology is the thread on which the tap screws onto the bucket.
The second part is the harness. The models themselves have converged. What has not converged is everything around them. Karpathy gave this layer its names — context engineering, then agentic engineering — and the industry compressed it into a formula: agent = model + harness. The harness is the control loop that decomposes the task, feeds the model exactly the right slice of the bucket at exactly the right moment, retries on failure, and gates what the agent is allowed to touch. One 2026 estimate holds that roughly 88% of AI agent projects never reach production, mostly because the harness is too fragile. Read that number the way I do: the scarce engineering is not the model. It is the plumbing between the bucket and the spout.
The third part is the pour. The last inch of the tap is presentation: how the squeezed, massaged water actually reaches the person. This is where generative UI is quietly becoming its own moat — interfaces that assemble themselves at runtime based on what the data and the moment demand, instead of forcing everything through a chat box. The same data, poured well, is worth multiples of the same data dumped in a dashboard. Anyone who has watched a partner's eyes glaze at a spreadsheet and then light up at a one-page founder-signal report knows the pour is not cosmetic. The pour is the product.
Jerry Chen saw the shape of this a long time ago. His Systems of Intelligence framework split enterprise software into systems of record, systems of engagement, and the AI layer between them that combines proprietary data into insight. In my vocabulary: the system of record is the bucket, and the system of intelligence plus the system of engagement — the squeeze plus the pour — is the tap. He called the middle layer the next defensible business in 2017. It took frontier models arriving for the tap to become buildable by a two-person team.
The scorecard
Half of this essay is an investment memo, so here is the memo. When I look at an AI company now, I score it on four questions, and a company needs all four.
Does it own the catchment? Not "has data" — sits upstream of a flow nobody else can stand in. Granola is in the room while the meeting happens. Cursor is in the loop while the code gets written. Set that against the argument that value accrues at the bottom of the stack while the application layer gets commoditized — which, read closely, is really an argument about companies that own neither the catchment nor the pour. A wrapper under the storm drain deserves every bit of that skepticism. A wrapper upstream of a proprietary flow is a different animal wearing the same fur.
Does it own the meaning? Data joined under someone else's ontology is someone else's asset. If the definitions live in the customer's head, the bucket walks out the door with the churn. If they live on a partner's platform, the bucket can be siphoned through an API. The ontology is the one part of the moat that never shows up in a metrics deck, which is exactly why it gets skipped — and exactly why Palantir gets to charge what it charges.
Does the tap feed the bucket? Every pour should come back as new captured data — the correction the user made, the recommendation they accepted, the decision they took. This is the flywheel investors actually underwrite now, and it is close to binary: a tap without return-flow is a services business, and a bucket without a tap is a lake.
Is the bill at the tap? Customers should pay for pressure at the point of use — the report, the answer, the action taken — never for storage. Storage pricing races to zero the moment a second vendor shows up. Pressure pricing compounds, because every month of flow makes the next pour better than anything a competitor can offer on day one.
Companies that clear all four are rarer than the funding environment implies, and almost none of them market themselves this way. They look like a note-taker, a dictation app, a hackathon platform. That is the tell, not the objection. The model layer is a physics competition between giants, and the bottom of the stack is a capex war. The durable middle class of AI value is BAT pairs — one bucket, one tap, screwed together, spinning.
Turning the tap on ourselves
I spent the spring building buckets. The hackathon build process — every prompt, every false start, every agent trace, every judge question — flows into infrastructure we control, and I wrote thousands of words about why that catchment matters. What I have underbuilt, and underwritten about, is the tap.
Because here is the uncomfortable audit: we have buckets filling every event, and most of that water has not been poured for anyone yet. The judge who spent a Saturday scoring twelve teams has not received the intelligence her scores generated. The partner who posed a challenge brief has not seen the pattern report across every team that attacked it. The builder who shipped through our stack has not gotten back the mirror of how they actually build. Water, sitting.
So the rule for the rest of this year is the BAT loop, applied weekly: every bucket we operate must have a named tap, a named drinker, and a return-flow. Founder-signal reports to partners. Feedback intelligence to judges. Build-pattern mirrors to builders. Each pour captured back into the bucket as new data. If a bucket has no tap after a quarter, we either build the tap or stop pretending the bucket is strategy.
The hyperscalers taught everyone the first move: build the bucket where the rain is. The next decade belongs to whoever masters the second move. Collect the rainwater, yes. Then squeeze it, shape it, and open the tap — and let it pour out gold.
Comments
Loading…