Who: Ram Kumar is a core contributor at OpenLedger, an AI blockchain platform that lets people contribute and get paid for the data sets that train AI models. He has spent close to a decade building at the intersection of blockchain and machine learning, including enterprise work with companies like Walmart, Sony, and GSK.
What: The interview covers why OpenLedger exists: as AI models become more specialized, they'll need data that only individuals and enterprises hold, and Kumar argues the people who contribute that data deserve a share of the value it creates.
Traction: OpenLedger has attracted close to a million users contributing data sets and about 10 ecosystem projects building AI models on top of the platform. Its enterprise track record includes work with Walmart, Sony, and GSK, and independent developers in the Asian region are already building specialized models, like a sleep-data model, using OpenLedger's contribution tools.
In this interview, Ram Kumar, a core contributor at OpenLedger, explains why data ownership will matter more as AI models become increasingly specialized. He describes OpenLedger's proof of attribution system, which tracks who contributed data and ensures they get paid when it's used, and the community-building work needed to keep decentralized systems from being absorbed by large centralized players. He closes with examples of niche AI models already being built on the platform and his broader concern that AI development stay responsible and rewarding for the people whose data powers it.
Key Takeaways
AI Data Contributors Should Share In The Value
Ram Kumar argues that people whose knowledge trains revenue-generating AI systems should receive part of the resulting value. OpenLedger’s premise is that data contribution can become an economic opportunity instead of allowing large organizations to capture all the benefits.
Specialized Knowledge Makes Better AI Possible
Kumar says internet data is generalized, while important knowledge often remains inside enterprises or in individual practice. A surgeon’s experience, an artist’s process, or a firm’s internal expertise may never become a blog, yet specialized AI will need exactly that knowledge.
Open Systems Need Communities With Shared Goals
Kumar believes decentralized infrastructure will only resist concentration if people build together. Strong culture, common goals, and systems that are open, verifiable, and rewarding can coordinate contributors around an alternative to large centralized firms.
Proof Of Attribution Makes Data Ownership Verifiable
OpenLedger’s proof of attribution records who contributed data, how it was used, and what payment is owed. Kumar presents this as a trust mechanism that lets contributors rely on the system and its code rather than simply believing a platform’s claims.
Unique Datasets Enable Narrow, Innovative Models
Kumar points to a sleep model built from high-quality data gathered across different regions and races. The example shows how specialized models can emerge from datasets that are both unusual and valuable, especially when paired with health information for a focused use case.
Responsible AI Depends On Recognizing Data Value
Kumar expects general models to give way to specialized agents for healthcare, law, and other fields. His forward-looking concern is that these systems remain responsible and that people understand their data’s importance before centralized ecosystems consume it without reward.
Watch the full interview now on EO's YouTube channel! Below is the complete transcription of the interview. Minor edits have been made for clarity and readability.
Introducing Ram Kumar, core contributor of OpenLedger
I'm Ram. I'm one of the core contributors at OpenLedger. OpenLedger is an AI blockchain where people have the data sets. AI systems need these data sets. So as an application, you can use OpenLedger to go ahead and contribute a data set that you own.
We have a lot of data contributors that initially came on board. We have close to about 10 ecosystem projects building AI models on us. We have close to about a million users who are contributing data sets for that.
OpenLedger and Data Ownership
I've been in this industry close to a decade right now. The idea was to build an R&D company around blockchain and machine learning. We saw there is a need for enterprises to bring in fairness and like transparency within their ecosystem like within their organization.
We had an opportunity to work with enterprises like Walmart, Sony, GSK and many more.
And what we realized is that especially a technology like blockchain brings equality among every user that uses that. OpenLedger is a contribution from that. The idea was to not just service enterprises but build a product that can be used by anyone across the globe and figure out how AI is impacting everyone's lives.
We all have this epiphany at one point in time where you have conversations with your friend about a product that you want to buy and you see that ad on Instagram. You know that your data is being used. I've had multiple epiphanies of that, and we've worked with firms where that is visible, right? People's information was used to make their product better. Sure, it gave convenience, right? But it also took privacy. That's going to happen with AI as well. AI is going to make money. All these large organizations are going to make money out of it, but you're not going to be part of that. As we evolve, right, as models evolve, models will become specialized where they need data sets from people. And in that case, we need to make sure that we can own our data and we get paid for it. And that's what OpenLedger is trying to solve. It's a platform where users can come and contribute data sets, which could be, let's say, a knowledge that they have about trading or a knowledge about a particular subject.
Let's say I know how to cook well. I can go ahead and contribute that, and then models can use this data. And if they use that data and they build an AI out of that and this AI makes revenue or creates an impact,you should be part of that. You should get a piece of that revenue.
We want to bring in farmers to this ecosystem. The people who contribute data or a model developer or a compute provider or any kind of resource provider gets paid as part of the process, and that's what OpenLedger is all about. A lot of people say that AI is going to make people lose jobs.
I don't think so. It might be a temporary thing, but it's going to create a lot of jobs. Data contribution itself could be a great gig economy. Data is also very relevant with enterprises. Data you find on the internet is very generalized, but the data that you would find in a firm, in an enterprise, is very specialized, right? And all of this knowledge did not come on the internet, right? People don't write blogs about it. A surgeon doesn't write about how, what his experience in actually doing the surgery. An artist doesn't write about how he actually painted a picture. So it all comes down to the individual person's knowledge that they own. This knowledge would be needed for AI, right, for AI to truly get into all parts of our lives, more than just a chatbot. It needs to know knowledge about the entire world.
It needs to know about very intricate details, right? In that case, people would reach out to individual users to get their knowledge.
So we knew that the ChatGPT moment was not just going to be a spark. It's going to become much bigger. We're going to build AI systems that are very specialized in various use cases. But there is not enough data out there on the internet. We've seen a lot of independent developers who have a lot of innovative ideas who are actually building interesting AI models in the Asian region. They contribute data sets by using our nodes, which is basically a node they can download and have it as a plug-in.
They can contribute data sets for that. If you take a look at Web3's nature, the internet was supposed to be decentralized, but because of convenience we let larger organizations take that.
So making sure that it does not go back to a bunch of centralized larger firms is very important.
We don't have money to fight for it. All we have is our own power, people coming together and building something against the larger organizations. In order for that to happen, you need to make people come together, and community building is very important to us, as part of that: building a very strong culture, building a very strong community.
Having the same goal, building systems that are open, verifiable, and rewarding is what makes people come together. Let's take an example of Ethereum itself. Ethereum is a very community-driven blockchain and that is why it's so strong today. Even though it has its ups and downs, Ethereum as an ecosystem is very strong.
That's why I think community is very important in what we're building. So to explain proof of attribution in a very simple manner, what if there's a tracking mechanism where you can see who actually contributed for all of this, and you can also see it on chain that everyone who contributed are getting paid for it.
So it's a tamper-proof record of your data's ownership. It's a record of how your data was used and it's also a record of how much you're going to get paid if your data was used as well. So the reason why we have the proof of attribution is because we need to have a trustless system where they don't have to believe Ram, right? They can believe the system, they can believe the code. I think that's very important. As a data contributor, you can go ahead and choose the model that needs the data and you can start contributing data sets to that. It's that cyclic ecosystem you want to build. All the data that is contributed is recorded on the blockchain so that we can track this and we can pay this, as a user. I know that I can prove my ownership by having it on chain. Another interesting side that we have started to see: smaller model developers and innovative people who want to build something interesting have started using our product.
A good example that's being built on OpenLedger is that a bunch of doctors are building a sleep model.
The model is trained on sleep data sets, high quality sleep data sets across the globe. They want to get access to the sleep data sets from various parts of the world so they can cover various races.
This particular data set is very unique because they're going to correlate that with their health vitals, and then once the model is ready you can just upload your sleep data and then it's just going to tell you what your body vital looks like. So like this, we have very interesting players who are building models on us, very niche, innovative ideas where you need data which is quite unique.
We are quite excited about them. So a lot of people ask me like why this has to use blockchain. It could be a traditional AI company, but I don't think so. If you take a look at generalized models, the era of that will slowly fade away.
Specialized AI Models
Agents will become much more specialized. There'll be an agent for healthcare. There'll be an agent for legal. There'll be agents for every other sector out there. And you can have a general model power that, but you need to have a specialized model that powers it. And the data set is among actually people. And these data sets can eventually become models, specialized AI models which then can be consumed by apps and agents that are going to be built on top of that. So I think every aspect of our life will change.
How we take a ride home, our doctor visits and what we learn from all of that will change. I learned the world through the internet. I think for my daughter, AI will create a huge impact. They will learn the world through AI. AI has to be responsible.
So building a responsible AI system is upon us. If we encourage larger ecosystems to go ahead and consume our data and not reward us, then that's what is going to happen.
Right? I think now it's time to change it. If we can realize how important our data is, probably that's the best output that we can see out of AI.
Join the 1.5M+ founders inbox to get the latest updates.