Rethinking the Web in the age of AI
I want to talk about the Web. Why now? Because AI has wandered into it.
Here is a brief history of the Web.
In 1989, Tim Berners-Lee proposed the World Wide Web: a system for sharing information through hypertext.
In the 1990s, the Web became widely accessible, and what we call Web1 emerged: a read-only Web of documents and websites.
In the 2000s, Web2 emerged. The Web became interactive and social: a read-and-write Web built around platforms.
In the 2010s, Web3 introduced another idea: a read-write-and-own Web, enabled by decentralized networks and programmable assets.
In the 2020s, AI entered the Web.
Ok, the history of web has always been centered on information for human. It was designed for humans to read, write, and navigate information.
Of course, machines have been processing the Web for decades. Search engine crawl, index, and rank Web pages. But ultimately it’s for human to effectively explore the Web.
Machines have also been collecting information from the Web for specific purposes: extracting product information, collecting financial data, and building specialized datasets.
But these systems generally know what information they are looking for.
In contrast, AI doesn’t know in necessarily know, in advance, which information it needs. AI uses the Web to learn about the world in general. The Web contains an enormous amount of information about people, places, events, science, culture, and almost everything else. By training on this information, AI accumulates knowledge about the world in its parameters effectively.
We have long imagined AI as something that could eventually know everything. Because we have mapped so much of our knowledge onto the Web, the Web has become fertile ground for AI. But it has a problem.
** The infrastructure is not built for this!**
Information on the Web is distributed across countless servers and websites, often far apart from each other. It is also represented in forms designed primarily for humans to read and navigate.
I want to provide an easier way for AI to access the ground truth.
We already have a lot of common knowledge that we more or less agree on.
If someone wants to publish information, they can create a byte that machines can easily understand.
I think AI can help with this process. It is kind of recursive. Humans provide information, AI helps structure it, and then humans can verify and publish it. Once it is published in a machine-readable form, other AI systems can directly use it.
Then, we need to store or bind this information to some kind of trust layer, where the person or organization who publishes the information can control it.
Concretely, I think we can use Ethereum as this trust layer. The knowledge itself can be managed as JSON or some other machine-readable format.
Depending on the size of the document, storing everything on Ethereum may be too expensive or simply unnecessary. So we could use some kind of decentralized storage for the actual information, and use Ethereum to manage the ownership and editing permissions.
This way, the information can be distributed, while each participant can still keep the whole information in their own storage.
If someone wants to make this information available for AI to learn from, they just need to run their own storage node.
AI can then access the information from these nodes, while the trust layer tells it where the information comes from and who controls it.
I think this solves the problem we had before.
I haven’t implemented this yet, and I haven’t really thought about what specific technologies I would use to build it.
Then a website becomes just a presentation layer, like Web1. It can be created from a subset of documents, depending on what we want to present to people.

And when AI wants to learn from the information, it does not need to go through the website. It could just run a storage node and access the documents directly.
The difficult part is not really building the system itself. It is getting the whole Web to adopt it. But this system doesn’t need to break or replace the existing Web. Websites can continue to work exactly as they do today. From the outside, it may even be difficult to tell whether a website is using this system or not.
This is just my imagination of the next age of the Web.