Dave: Hey everyone, Dave here with another episode of the Philly Tech Connect podcast. Today I’m speaking with Michael Kener, a 3X Silicon Valley unicorn veteran with IPO experience. He co-founded Data Squirrel, which is an AI-native data engineering platform and neuro-symbolic AI engine. He’s writing a book about data products and Quantum Computing, so there’s a lot of terminology here that I’d love to dive into and learn more about. Michael, how are you doing today?
Michael: Fantastic, thanks for having me. It’s a pleasure.
Dave: Tell us a little bit about Data Squirrel and your reason for getting that started, and then let’s see where that takes us in terms of demystifying some of the words in your bio.
Michael: Sure. The last 20-something years of my career has been very data-focused, even though I’ve worked very hands-on at times with many technologies. My career has always revolved around data, and over the last 10 years, specifically when I’ve spent time in Silicon Valley startups and unicorns, which, if you don’t know what a unicorn is, it’s a company that’s valued at a billion dollars or more. I was working on some very cutting-edge things in terms of technology, and everyone I worked with had a master’s in computer science or electrical engineering, and some of them were so tough that it really required that. So, the reason why we started Data Squirrel—my co-founder actually has a Ph.D. in database systems. I myself have advanced degrees; not only do I have my MBA from a local school, Villanova, but I did graduate work at Harvard in artificial intelligence and data science, and MIT in Quantum Computing. We looked at the data technology landscapes, and we’re just like, this is way too hard, and not everyone has the time or the mental capacity or focus to go to a Harvard or do a Ph.D., and because of that, people are not getting the data that they need. If you talk to people and you look at the amount of data out there, it’s apparent people are basically drowning in data, and one of the top complaints I hear is, “I have the data, I just can’t get to it in an efficient manner.” So that’s why we started Data Squirrel. Myself and my co-founder, both of us have years of experience in Silicon Valley technology, and so on. That’s why we really started it.
Dave: So you’re saying you can’t get to the data in an efficient manner. Now, does that mean like accessing the data in terms of pulling it from the database and putting it into a spreadsheet, or whatever it might be, or more the analysis of the data, interpreting it?
Michael: It’s really both. So I’ll give you a really good example. If you swipe your credit card at the gas station, what do you think happens?
Dave: So I would assume that my number that’s attached to the credit card goes into their sort of like their Point of Sales system, and they send some message to the credit card company that I owe a certain amount of money, and then the credit card company also gets some data that says, “Hey, I was at a gas station, and I paid $40 to fill up,” something like that.
Michael: Close. What actually happens is that message is sent through something like Apache Kafka, which is an event-based or event streaming technology. And yes, it does check your balance to say, “Hey, can David actually charge this amount of money?” Let’s just say it’s a preset amount, $100; sometimes you see $100 get charged. It’s really a hold, but what actually happens in addition to that, and this is something I’ve actually worked on personally, is that message gets sent off to about 40 other different algorithms that are doing various things, and then the message gets returned back to the gas pump saying, “Yes, David can charge this money.” That’s the kind of complexity I’m talking about. Yeah, you think it’s just a point of sales system, but when you think about how many transactions are hitting a credit card company at any given moment in time, the scale is kind of mind-boggling to a lot of people. And is there a need for it to be that complex for there to be 40 different algorithms, or is this more just kind of this technical sort of tumbleweed that has over the years kind of become what it is?
Dave: Yeah, the best way I could describe it is a stalk of corn. Okay, and whether you like corn or not is sure, I’m fine with it.
Michael: Yeah, there are pieces around the like an actual husk of corn that you have to pull off to get to the actual corn to you for you can eat, right? And is it necessary for you to pull it off? Yeah, but like that stuff kind of just ends up growing on top of it, right? So to answer your question, somewhere in between really good companies with really great data strategies understand their there’s a function or something there that they need to have, and then somewhere along the line, some lines get blurry, and people just start adding on top of it these layers that need to be peeled in order to actually like eat and enjoy the corn.
Dave: Okay, so neuro-symbolic, is this related to what we have been speaking about? Where does that fit in? It’s the first time I’ve ever heard that term.
Michael: So neural symbolic, it sounds like a really hyper-complex term. It’s actually not. You’ll, I think you’ll immediately understand what it really is. Neural symbolic, neural is the AI piece, symbolic is the data piece, and all Data Squirrel really does is act like a bridge between the two. So if you think about OpenAI or DeepSeek, we’re hearing a lot about DeepSeek right now in the news. If you talk to anyone who has any sense of AI or data, one of the big problems they will tell you, and I’ve seen this firsthand because I’ve run rather large data science teams, is AI needs the right data for it to be impactful. Have you ever heard the term hallucination in AI?
Dave: Let’s assume that most people have not and define it.
Michael: It’s a lie. That’s what it is. The AI is making something up because it doesn’t have the right answer. It doesn’t know it doesn’t have the right answer; it’s just trying to serve a response to the question. It’s kind of like talking to an 8-year-old, and you’re asking them what quantum physics is, and they just start making up an answer. Well, having run data science teams in the past, I know that data scientists often spend 80 to 90% of their time being data engineers, essentially plumbers, trying to get the right data to their models so the models could be accurate and provide the right answers to minimize hallucinations because, I mean, this is a major problem. You know, maybe not, you know, hey for me, kind of trying to be like, “Oh, how many people live in Paris?” or you know some type of general query, not such a big deal, but when you know lives are on the line if it’s you know cybersecurity or or legal or something like that to be given just kind of a non-answer answer is a big problem, I assume.
Michael: Yeah, yeah, and it’s gonna get worse over time. And you know, I guess the AI models are not, um, you said they’re not aware that they don’t know the answer, so they’re not able to say, “Hey, I don’t have an answer for this query.”
Dave: Yeah, they have the makings of an idea what might be, but and they go down that line of reasoning and the chain of thought or co uh that they respond to it, but they don’t have the full picture. And what you’ll see over time, by the way, in the next few years, is you have large language models, and you’re going to see smaller language models, large language models that are hyper-specialized for a very specific field. I’ve been saying this now for a few years. I think in the next five to seven years, if you’re a doctor or a lawyer, you might get sued because you don’t use a hyper-specialized large language model or AI. Uh, so neural symbolic is just getting the AI the right data and combining it.
Dave: Yeah, I’ve seen come across some articles recently where um some professor or or legal individual um kind of got caught um for providing something that was not factual, not you know, not didn’t have the right citations because they were leveraging AI. Is that you know kind of related to what you’re speaking?
Michael: Yeah, matter of fact, I think it was like two years ago that a few lawyers were disbarred because they used OpenAI to prep a legal brief, if I remember correctly, and they were disbarred at the time. Moving forward, I predict the large language models and AI will get so much better, and you’re seeing it already that when you go to the doctor, your doctor is now going to have, hopefully, a tool to say, “Hey, uh, I’m seeing Dave, uh, give me a rundown of his family history. What are potential issues that he might encounter in his health based on genetic markers or just family history?” And I’ll give you another example, cancer is rampant in my family, so we are very vigilant uh around that, and in the future, AI may come out with some preventative medicine or even God willing some cures, and my doctor who I just went to see yesterday for my yearly physical, in five-six years, may be able to see like, “Oh, Mike, think about doing X, Y, and Z because our latest research tells us uh that and the latest AI discoveries tell us that will help you prevent colon cancer.” So yeah.
Dave: So I’m curious when you talk about this, you know, increasing specificity of LLMS, do you imagine that as being still kind of under the umbrella of the way we think of LLMS now? So for example, like imagine you know, OpenAI.legal, OpenMed, or that each kind of private company has its own LLMS that is you know, supported by its own data that it owns?
Michael: I think what you’re gonna see is the layer effect. So you have an LLM like OpenAI, right? Well, out of what may rise to the top of that LLM is a specialized LLM in the legal or medical or construction field, right, or material engineering field. So you’ll have a specialized LLM, and then you’ll have a hyper-specialized LLM that sits on top of that, where the hyper-specialized LLM may just be around cancer research or trust and wills or concrete, right? Uh, as so as it may sound, as you go down the line, you LLMs or I should say as you go through the evolution of AI, what you’re going to get to is not only this large umbrella of knowledge that an LLM captures but the reasoning behind it. That’s what an AI agent really is. It could take a line of reasoning, execute tasks against the line of reasoning. So you hear AI agent; that’s what that’s all it is. It’s something that can say, “Oh, um, it’s the 15th of the month, and based on Dave’s schedule, I see that he typically orders his protein shakes this time a month. I’m just G to go ahead and do that unless he tells me not to,” and it’s going to go ahead and do that.
Dave: No, it makes perfect sense. Um, yeah, you I would assume that in AI LM model that is specific to a particular you know industry and trained almost exclusively on that data would be more accurate and more powerful in that space. Um, so it seems like a natural evolution. Um, I guess we’ll see. You mentioned you’re writing a book on data products. I’m just curious, um, kind of what you know how that fits into everything else you’re working on.
Michael: Yeah, so a data product is really an interesting way of reimagining how we do integration. Uh, there was an article in Harvard Business Review written by McKinsey Consulting, the famous big consulting firm, uh, around how and it makes complete sense to me because I’ve been in it forever. Uh, a lot of times when you go into a company, they’re like, “Oh, we have to integrate this data.” No one’s really looking at the data and saying, “Hey, how are we producing the data and are we producing it correctly?” So you hear these terms, “data is a product” or “data product,” and I like to think of it almost like Air Jordans. Now, you look young; I’m just going to assume you’re young. Is I am not debatable, but okay, I remember when Air Jordans came out, and if you think about an Air Jordan or sneaker, any you know clothing or food that you get, when Nike produced the first Air Jordan, they had to create the facility, the manufacturing line, the supply chain, uh, the marketing around it, but then they hit generation two. Do you think they went out and just rebuilt another factory and rebuilt another supply chain? No, they just reused; they modified it a bit but they reused it, and they produced the sneaker pretty much the same way, a little bit changing design, for instance, and maybe some of the materials, but the process is the same. In technology data, every time we have one of these projects that we have to integrate data, we kind of just rebuild the world, uh, in some sense. You in some ways, good companies understand this, and they have a really good process for a project but not really a good project for or a good methodology for the process of producing data and maintaining data. So you know why you, I have a visual of this, uh, the data product framework. The very base is, “What’s the data product you’re trying to produce? What questions is that data answering?” Then there’s the technical piece, “How you like, where’s the sources, where the attributes,” and so on. Then there’s the iteration piece, the or the versioning, version one of the customer data product. Okay, who governs it? How often is it reviewed? And then at the very top is the value. What value does it provide to an organization? And when I say value, it’s often a very soft or nebulous term. I’ll give you a very good way to think about value, and it comes down to four things. Value is either you’re making money, you’re saving money, you’re protecting money, or some combination thereof. So and define that value could be is easy as, “Hey, it costs this much to store the data.” That’s one aspect of the value of the data, or that it supports this service that brings in x amount of dollars to the company every year or the organization. So uh, the whole idea behind the book in the data product book is to treat data and the process of making it like you’re almost manufacturing it, so it has a purpose. It’s not just data out there just to have data. It’s mindfully and thoughtfully looking at how you produce data and what the Harvard and McKinsey study has shown is if you do it that way, your cost of integration drops by between 80 to 90%. Nice. Sounds very pertinent to you know, the discussion that we’ve just had around LLMS and kind of the barrier to entry to building them and making them accurate is data, and so um, I’m definitely curious to see how that evolves.
Dave: Michael, for people that want to get in touch with you, learn more about Data Squirrel, the book you’re working on, anything else, um, how should they do so?
Michael: Uh, a few ways. Uh, you can always find me on LinkedIn. I go by cans, C-A-N-Z. I’ve been can since the first day of first grade. Uh, so you can find me there. Uh, you can also find this at Data Squirrel. Uh, data D-A-T-A Squirrel is spelled S-Q-R-L Doom. Uh, so you can find us those two places, and uh, when all else fails, uh, just look for me on cans. There’s only about four Michael Kener Aries in the world. Uh, one’s a brigadier general, I believe. Okay, good. St one’s a lawyer, and one’s another software engineer in, I believe, the Buffalo, New York area. Wow, I might name my next child Michael Kener just by, you know, feels like a winner.
Dave: Yeah, so there you go. Thanks for being out, Michael, appreciate it.
Michael: Appreciate it. Have a great day.