AlphaFold

Did you know an AI model won the 2024 Nobel Prize in Chemistry? Well, the model didn’t actually win but the work of two Google DeepMind researchers, Demis Hassabis and John Jumper along with David Baker (University of Washington) won it for their protein structure prediction work using the model.

AlphaFold is an AI model tailored specifically to understanding complex protein structures. The speed and scale AlphaFold provides to researchers unlocks accelerated drug discovery and deeper understanding of diseases. It’s truly one of the positive breakthroughs in AI that doesn’t get enough popular press.

An interesting article in the journal Nature this week highlights that AlphaFold is getting outperformed by pharmaceutical company models using proprietary data. AlphaFold was trained on public data. There are numerous competitors in the space across academic and commercial sectors. It’s also notable that AI models are configured and trained for different purposes. Some for speed, others for specificity. The point is that models can outperform across different objectives and domains.

However, here is another example of value in AI not accruing to the model layer, as I discussed last week. In this case, it highlights where proprietary data is worth more than the model.

Large companies are already evaluating the trade-offs when using external models from Anthropic or OpenAI with their proprietary data. Corporations fear the model providers can use client data to train models or to provide competing products. It’s not simply the data, but also the processes. Often proprietary processes or methods are just as valuable intellectual property as data. Can Anthropic incorporate the skill into CoWork or another application and resell it?

What about your small businesses or startup? AI has unleashed amazing productivity for less capitalized businesses to build products, replace or enhance labor, and produce other efficiencies. How should you think about sharing your data and knowledge with model vendors? Are you at risk of replacement through knowledge acquisition?

What To Do About Your Data

Of course, the answer is not to stop using AI.

AI is already so useful, you and your business can’t sit it out.

The question becomes what information you are willing to expose, under what terms, and in exchange for what benefit?

For most businesses, I think about this in three buckets: Share, Shield, and Isolate.

Share

For work that is already effectively public or commoditized, use away. Even at the free and consumer tiers. Examples include; rewriting marketing copy, brainstorming headlines, summarizing public research, drafting generic job descriptions, or explaining an Excel formula.

It is the cheapest intelligence you can buy, which usually means you are making a trade somewhere else. Depending on the product and your settings, your conversations may help improve the provider’s models or products.

If the information would not meaningfully hurt you if it appeared in a public document tomorrow, this is usually the lowest-cost place to use AI.

Shield

This is where most real business work should probably happen.

Paid business products and APIs increasingly come with contractual or default protections against using your inputs to train the underlying model. You are paying more for the intelligence, but part of what you are buying is a better boundary around your information.

That makes this the right layer for internal financial analysis, customer support workflows, operating documents, code, sales data, and other information that is sensitive but not truly existential to the business.

Your data is not the only thing you are exposing. You are also revealing how you work.

The prompts you write, the workflows you build, the sequence of tasks you automate, and the problems you repeatedly ask the system to solve are all signals about your business. Even if the raw data never becomes part of a model’s weights, your usage can still teach the provider what customers value and which products to build next.

For most companies, that is still a perfectly reasonable trade.

But it is worth recognizing that the thing you are protecting may be more than a spreadsheet with valuable data.

Isolate

Then there is the information you should treat differently.

This is where local models, self-hosted open-source systems, private infrastructure, or even air-gapped environments begin to make sense.

The standard should be simple:

If a competitor gained access to the distribution of this information, or learned exactly how you use it, would it materially weaken your advantage?

Examples include proprietary research, unique underwriting datasets, or customer behavior that nobody else has, or highly differentiated internal workflow. Or a dataset that becomes more valuable precisely because nobody else can train on it.

If you don’t believe unique training data is valuable, Anthropic and other frontier labs have been hoarding rare books to feed into their models.

That is the kind of information you may want to isolate.

The tradeoff is that you give something up. The best external models benefit from enormous scale, constant improvement, and network effects. Running your own model usually means accepting more cost, more complexity, or somewhat less capability.

Sometimes that is worth it.

The mistake is treating every piece of data the same.

You do not need a fortress around your lunch order. And you probably should not paste your company’s most valuable proprietary dataset into the cheapest chatbot you can find.

The goal is not maximum privacy but to understand and be deliberate about what you are trading away in exchange for intelligence.

The Next Step

There is one final wrinkle for small businesses and startups.

Your individual data may not be especially valuable.

One accounting firm’s workflows, client questions, pricing decisions, and internal processes probably do not matter much to OpenAI, Anthropic, or Google.

But hundreds or thousands of accounting firms using the same systems might.

Taken together, those interactions can reveal how an industry actually works: which tasks consume the most time, which questions customers repeatedly ask, where judgment is required, which processes are standardized, and which parts of the job are easiest to automate.

The proprietary asset may not exist at the company level.

It may exist at the sector level.

That creates an uncomfortable dynamic. Each individual business can make a perfectly rational decision to use AI because the productivity gains are enormous. Collectively, though, an industry may be teaching the model providers how to perform more and more of the work that industry currently gets paid to do.

I still do not think the answer is to opt out.

The economics are too compelling, particularly for small businesses that can suddenly access capabilities that once required much larger teams and budgets.

But we should stop treating every interaction with AI as free.

Before handing over data, processes, or workflows, ask two questions:

How valuable is this intelligence to me?

And:

How valuable could what I am giving back become to someone else?

Most of the time, the trade will still be worth making.

Just make sure you know which side of the trade you are on.

My goal with The Leap is to provide you each Saturday with the knowledge, tools and lessons learned to help you get started and keep going toward building your future. 

Whether you are making the leap to startups, solo-entrepreneurship, freelancing, side hustles or other creative ventures, the tools and strategies to succeed in each are similar.