Dave: Hey everyone, Dave here with another episode of the Philly Tech Connect podcast. Today, I’m speaking with the founders of Setm, a company that aims to connect data silos. On my right is Jackson Morgan, the CEO and co-founder of Setm, where he leads engineering efforts to develop solutions for data interoperability and precision in large language models. Jackson has built carbon tracking systems for the Government of Scotland, supply chain solutions for the Internet of Production Alliance, and decentralized credential systems for MIT’s Digital Credentials Consortium, all while advancing open-source tools for the linked data ecosystem. Further to my right is Ian, the COO of Setm. He’s an MBA candidate at the University of Pennsylvania Wharton School of Business and has previously worked at the US Department of State and also Amazon. Guys, how are you doing?
Jackson: Doing fantastic, how are you doing, David?
Dave: I am doing well, thank you for joining the show. Very interesting bios, backgrounds, and sort of project product concept that you guys are working on. Tell me a little bit about how you each got into the space of, I guess I’ll call it decentralized tech, and I apologize if I say anything incorrect because it’s not my field. But tell me a little bit about how you got to where you are today and how you guys met each other.
Jackson: Yeah, cool, I guess I’ll start. Way back, pre-pandemic times, I always had this idea in college that the way that we do data is honestly really bad. The fact that data exists in data silos really limits your ability to use data. In order to connect data silos, you have to do all of this integration work, and it’s really annoying. So in college, I was doing some special projects to try to fix this and then I found out about this world called the linked data world, which proposes a solution to data silos, a way to connect data together so that you can utilize insights that don’t exist. At that time, I quit the job that I had coming out of college and decided I’m just going to apply to a bunch of academic conferences about linked data. In going to those academic conferences, if you’re familiar with the academic world, there’s something called posters. Posters basically get accepted to any conference that you give them to, so I created a poster, got into a few conferences, and at one conference, there’s this older gentleman who came up to me and asked where is the location of this event. I said I’m going to the same room, follow me, and when he came, when we went into the room, we sat down together and then the presenters kept speaking to him with deference. I wondered who is this person, and I eventually found out that this was Berners-Lee, the inventor of the worldwide web, who is a huge name in linked data. This goal that we are going to take all the data on the web and turn it into one unified database. So I met him then, he forgot about me, but I was able to network with a lot of his colleagues and eventually went to work for his company, Inrupt, building these linked data solutions, and eventually left Inrupt and started to build a consulting company around linked data, and I met Ian a few years ago. Ian has a real great business background to complement the technical background that I’m working with. So Ian, what’s your background?
Ian: I talk more about my background too and sort of how I came into contact with Jackson. So I met Jackson in I believe the summer of 2022 in Washington DC when I was living and working there. Although I hadn’t been exposed to Solid or Linked Data before, I was able to learn about it from him and when we first talked about it, it really stuck with me because I was working at the US Department of State and US state, like most government agencies, like most large established enterprises, they’re built on decades and decades of legacy IT infrastructure where you have all these different data silos across an organization. It’s really expensive to share that data as a line-level employee, it’s really frustrating because you know different data is located in different locations, you might not have access to it in a certain amount of time and you end up creating work products that you know are suboptimal even though all the resources are there. It’s simply just an administrative and technical barrier that’s preventing you from getting what needs to get out and what needs to get done. So, I was just very excited to learn there was a possible solution and I eventually left the Department of State and I went to the Wharton School of Business here in Philadelphia where serendipitously Jackson also had moved to Philadelphia. The key moment was I was in my valuations class with Professor David Wessel and he was giving us this story about how data is very important for large organizations and it’s like a moat right in businesses. You want moats, you want established things to keep your business able to compete against its rivals and its competitors but the irony with data is that this moat actually hurts you and it’s a moat that you dug yourself. So all these large companies like Oracle, SAP, AWS literally benefit from the administrative and technical inconsistencies you have within your organization and then realizing that what we can do at Setm, we actually transcend all those issues and all those problems and suddenly what can become one of the most expensive line items from an IT standpoint for a large established enterprise, we can reduce the costs drastically and then just even allow even more functionality for the users who deal with this data every single day.
Dave: Cool, sounds great. I have a lot of questions, hopefully they’re not like loaded questions but I’m curious, it’s something I don’t know that much about, but it’s like I feel like it’s a hot topic these days. So data silos, again the way I’m interpreting this is like a business, like a Facebook, like they have data on me as a user, who my friends are, what pages I visit, and that’s something that’s data that they own and I don’t necessarily know about. I don’t have access to it, and neither does LinkedIn or some other social media. And what I’m understanding is you know if we were to decentralize this, this data would sort of be out there for everybody and presumably, I would own it in some way. So I guess where I’m having difficulty understanding is, is number one, if the data is decentralized, you know why is that that make it that all of a sudden I own my data, you know isn’t that just actually everybody owns my data and do I want everybody to like see my data, you know what I’m saying?
Jackson: So this is the funny thing. So you are familiar with web 1.0, Web 2.0, Web 3.0 right? And a lot of times these days people talk about Web 3.0 and what they mean is they mean blockchain or they maybe they’re talking about distributed hash maps. But there actually was back in the 2000s a different Web 3.0 and this is the kind of web that we’re talking about, the web of linked data. So erase everything from your mind about blockchain about distributed hash maps, that’s not something that’s effective for the modern business world, it’s not something that I think is actually effective for individuals beyond just making a bunch of money through pump and dump schemes. The way the original Web 3.0 works is we’re going back to the roots everyone can choose what server they’re putting their data on, you’re not necessarily putting it on Facebook server, you’re not necessarily putting it on LinkedIn server, it’s your server or perhaps a trusted third party that you decide but it doesn’t matter what the data is, the data storage becomes a utility. It’s not bound to the service so let’s say let’s use a business example beyond a social media example. Healthcare data, there are large silo databases that contain your healthcare data, they might be on Epic Systems, they might be on Cerner systems these are all big silo databases that contain healthcare data and you don’t necessarily own your own healthcare data it’s not all in one place. But if we can desilo things if each healthcare provider could have a system that they host that brings in data from all of the various legacy systems together into one standardized place now you can perform queries across that data that can create metrics that you weren’t able to get before. So we’re not taking data and putting it out into the ether onto a blockchain, you’re still it’s a still a traditional method of storage on a centralized system it’s just that those systems have specifications in the same way that the worldwide web has a specification for every web browser to be compatible with every web page, these storage systems have specifications to make sure that they’re compatible with all the other storage systems that exist thus bringing all the data into one place, a place that you chose.
Dave: Okay, again I’m gonna continue with a line of potentially stupid questions but this just sort of how it has to be, if the data is centralized now and I’m sort of choosing you know maybe who’s hosting it where it’s hosted the server things like that who’s paying for this like is it me that I pay for it now because now it’s it I own it and the burden is it or who is who is in charge of the centralized data system whose responsibility is it to maintain it and pay and things like that?
Jackson: So one of the big problems that we’re addressing isn’t necessarily that we’re addressing first is a problem that is experienced by large enterprises so instead of you saying uh who owns the data is it me um let’s think about who owns the data at say a hospital or a finance institution. It’s still the hospital or Finance institution they have budgets IT budgets but their IT budgets are ballooned by the fact that there are a bunch of disparate data sources some might be on Oracle some might be on SAP and now you need to build a bunch of integrations together you can still that you can devote that IT budget towards having a linked data store in instead uh where now all of your legacy systems can come together inside of a a pool of data and uh you can now perform queries across it and on top of that that pool of data isn’t just a generic pool of data it’s it’s not like U say a um say say like a snowflake kind of a repository for data it is a pool of data that has access control so now different departments inside of your large organization are able to set access control for data but let you query that data as if it was abstractly one database let me make that a little bit less abstract uh let’s say I have uh different government departments government of veterans uh the Department of Veterans Affairs uh Department of Defense and uh uh and DHS uh or sorry yeah department of homeline security or Department of Health and Human Services let’s say Department of Health and Human Services uh these all have medical data that exists but all of this medical data is currently siloed now the Department of Defense doesn’t want to share absolutely everything with uh the Department of Veterans Affairs it only wants to share a portion of it if you put uh Medical Data onto your linked data store if the Department of Defense puts Medical Data onto its link data store and the VA puts Medical Data onto its link data store they can control access control rules but because the data is linked now the if the Department of Veterans Affairs wants to find out uh data Healthcare data that’s owned by the Department of Defense they can just perform a query to get all of the data that they have access to the data that they don’t have access to the data that the Department of Defense has prevented them from having access to isn’t provided in this query but it all works seamlessly from a developer perspective you don’t need to millions of dollars building Integrations between things it makes it so that the decision to share data is simply a policy decision not a engineering an engineering decision and if I can add one more thing too Dave um just to talk about the costs like for a large established Enterprise your data storage costs are probably like 200 to 300 million dollar a year there is a reason that Oracle is a $450 billion market cap app company uh we really want to change that business model because we think it is taking money away from organizations which can use that Capital much more effectively and what’s also really exciting about these linked data stores is that once you make those changes in those original systems that are linked to those link data stores that data then is also switched in everything else that’s also connected to those link data stores so all this data now changes in real time at the end user level without them having to be educated on any new system any new procedure so you can keep all these Legacy it systems whilst having this sort of data translator tool exist on top of your IT infrastructure.
Dave: Okay, our goal is we we we really don’t want to screw you over from a data perspective um there there is a a story I have a friend who worked for Amazon um and Bezos and Ellison uh of course Jeff Bezos who was at the time CEO of Amazon isn’t currently CEO of Amazon and Larry Ellison the CEO of Oracle got into a public spat on CNBC uh David David Ellison said that uh Amazon will not be able to migrate off of Oracle system they had a very antagonistic uh relationship with each other and uh Oracle just said hey our systems are too difficult to migrate off of that’s a feature that’s why Oracle is so uh is so well owed um Bezos took that as a challenge and then devoted armies of Engineers just to migrating off of uh Oracle systems I had a friend who was a part of that army of Engineers it was the the the worst time in her life it was very difficult to do um but we want to provide solutions that makes it easy for organizations to integrate their various data sets um and and not to just uh basically take ownership over their data over data that those organizations should own themselves so if I understand correctly um centralizing the data is kind of done like at the level of an organization it’s not like the worldwide web where sort of you’re either on it or you’re not and and everybody’s kind of hanging out with each other it’s you know within a particular organization of the government they have data and all these different silos and they want to be able to centralize it so that different departments people can access it more like seamlessly or they can to share with other organizations that maybe have a similar Mission something like that.
Jackson: Yes uh so this is this is the the service that set meld is focused on um from a technical perspective uh the open source work that we’ve done in the past and are continuing to do is focused on uh everything so you as an individual could also deploy your own uh linked data store and we’re providing open source tools to work with that in fact the the call that I just got off of was talking with some of my friends in Europe working on open source tools that work for everyone um but with set meld we are laser focused on Enterprise customers that really want to uh own their own data and integrate their own data and to help them get out of a the sort of precarious data ownership ecosystem that exists today and if if I saw you guys are employing solid which appears to be sort of a data architecture library that uh will uh support this centralized um format for a data ecosystem how does solid play into this and what is solid so solid is a set of standards that is created by uh the by an international audience of people uh it’s you can think of it in the same way that the worldwide web is a set of Standards so if you are uh working with setm you don’t need to understand anything about solid this is the the technology that is underneath the Hood um that makes everything work it’s the technology that powers linked data Stores um and the the reason why it’s important to have a standard is because now if I have a bunch of different departments that are uh trying to share data with each other that’s not possible unless everything is standardized so we employed the solid standard which is an open- Source standard means that you’re not going to be locked in to a uh to a proprietary standard for your data finally um and uh and it’s a standard that has been developed over the the course of almost a decade at this point uh and has had a lot of its uh has had a lot of its use cases thought through by and Army of people across industry and Academia so we find it to be a very reliable standard uh for delivering data interoperability for our customers very cool um tell me a little bit on the business side of things Ian you know um the product that you guys are are working on um is there a you know where are we at do we have an upcoming launch date and and how does like the business model work for something like this is this your traditional SAS or is it like uh based on the amount of data that you store you you pay a tiered amount how are you guys thinking about approaching this.
Ian: Yeah, so the way we’re thinking about approaching it is really partnering with large established enterprises. We really want to get in touch with CIOs, senior VPs of engineering who are willing to really look at a novel solution to this systematic issue. We’ve really been talking with a lot of different stakeholders, we had a briefing at the Pentagon recently, I talked to a CEO of a large Pennsylvania based healthcare organization about what we could do too. But right now we’re still looking for customers, at the same time we’re actually participating in a University of Pennsylvania Wharton and then Perlman School of Medicine joint health tech accelerator which will allow us to work more with the Philadelphia based medical community and really hone in on a specific use case. We’ve been thinking and focusing on immunology records and allergy records as a possible use case but we’d really love to think about something where here’s a very specific set of data stores we know that we can integrate, we know we can connect, and then from that building out an entire IT infrastructure platform.
Dave: Sounds really innovative. Excited to kind of follow along with the journey and hear where you guys end up over the next months and so on. For others who want to learn more about Setm and you know what you guys are up to, how should they get in touch?
Jackson: You can go to set.com, that’s sldd.com, contact us there. You can also reach out to us, either of our emails, I’m jackson@set.com, so if you want to contact me you can do it there.
Ian: And I’m ian@set.com, so very easy to remember.
Dave: Yes, perfect. Thanks so much for your time today, guys.
Jackson & Ian: Thank you, Dave, absolutely.