Jerry Zhang Podcast Transcript
Jerry Zhang joins host Brian Thomas on The Digital Executive Podcast.
Brian Thomas: Welcome to The Digital Executive. Today’s guest is Jerry Zhang. Lemma is a production monitoring observability platform for AI agents, founded by Jerry Zhang and Cole Gawin, who met as freshmen in USC’s startup incubator. Lemma helps engineering teams uncover hidden semantic failures, diagnose root causes across thousands of traces, and deploy automated fixes, enabling AI agents to continuously improve from real-world usage instead of silently degrading.
To learn more, visit www.uselemma.ai. Well, good afternoon, Jerry. Welcome to the show
Jerry Zhang: Hey, Brian, how you doing? Thanks for having me.
Brian Thomas: Absolutely, my friend. I appreciate it. You’re hailing out of San Francisco, the Bay Area. I’m in Kansas City. Two-hour difference, but I know it takes a lot of work to get the calendars to sync and time zones, so I really appreciate you.
And Jerry, if you don’t mind, I’m gonna jump right into your first question. You and Cole met as freshmen at USC Startup Incubator and then went off to work on AI systems at separate AI native startups. Cole was building agents at health– in healthcare at Tandem, and then you were working on chip design agents at Chipstack.
Before you guys reunited back together and, and found Le- Lemma, how did you get to know each other and start this work off as co-founders?
Jerry Zhang: Yeah would love to answer that question. First of all, Brian, thank you so much for having me on the podcast today. I can totally tell you did your research and then looked into exactly how me and Cole met.
It’s just like you mentioned both me and Cole were co-founders. Originally, we’ve been best friends since freshman year, right? So one of the first times we actually met each other was in the dorms, just casually hanging out and stuff like that. And then one of the first places where we first started getting close was this on-campus incubator called SCP.
And then later, we worked in Lava as well, and a bunch of these different college clubs where we were super interested in entrepreneurship and just building a bunch of really cool projects together. And one of the first projects that we actually worked on was this company called Clinicode, which was like an AI healthcare company that was doing a lot of the medical coding and billing automations and things like that.
So both me and Cole were really passionate about sort of building our own things, taking things from zero to one and that’s one of the first instances where we realized that we actually worked really well together, right? I think what makes me and Cole really special about being co-founders is the fact that I’m personally really passionate about talking to people getting to know a bunch of different people but at– on the other side, like Cole is so deeply technical and an expert in what he does and a lot of the research and the more technical side of things that we really make a great co-founder pair, right?
So I think it was really that first experience working together on a real company especially solving a problem that we were really passionate about in healthcare, that made us realize that we were so passionate about working together and working on startups and building something together. Then as we were working on the project, it kind of began to lost steam because we were doing school and then also the startup that we were working on at the same times.
And then we were also just freshmen at college, right? So we wanted to gain a little bit more experience before we kind of like worked on our next endeavor or focused on anything else. So we each went off to like sort of build our own expertise in some different area, but then we somehow ended up coming back and converging together as well.
So I was working as a software engineer at this company called Chipstack, where we were doing AI agents for chip design that ended up getting acquired by Cadence. Cole was building agents as an ML engineer at this company called Tandem which is now called Forest. Now they’re like a billion-dollar healthcare unicorn and stuff like that.
And then we came together and realized one summer while we were chatting that at both of our internships and where we were working that we were facing the exact same problem of like agent systems and like building stuff and realized that this was a really pressing pain and wanted to work together and come build a product within the space.
Brian Thomas: That’s awesome. I love the backstory. That’s just– That’s what always makes the podcast interesting is, you know, everybody loves a good story. But you both hit it off as freshmen, that– which is awesome. Both interested in entrepreneurship. You had a lot of passion there. You worked together obviously at USC Startup Incubator, and then, of course, that s– friendship of yours grew, got stronger.
But you also highlighted or, or noted that your skill sets also complement each other. And after you went off and worked at your own separate companies and had that work experience you came back together and saw that there was that need or that gap in the market, and that’s kinda what we’re gonna get into next.
So, Jerry, you said you started Lemma because you experienced the pain of building AI agents firsthand, agents that would appear to work but weren’t reliable enough in production. Can you take us into a specific moment where that pain became unbearable and how that frustration crystallized into this idea of Lemma?
Jerry Zhang: Yeah, absolutely. And I think what you said was absolutely right. I think as co-founders, you wanna be, like, kinda like friends before you jump into business together ’cause that’s how we put up with everything and, like, sometimes we get annoyed at each other. But at the end of the day we just really enjoy building it with each other and things like that.
But yeah, I think credit, credit goes to Cole as, like, being the first person to kind of experience that pain point. And honestly, like, he was the one that came up with the idea for Lemma and really, like, was, like, the first person to, like, bring this up and stuff like that. But based off of what he said I remember, like, one summer when we were just, like, calling, he was telling me how, as an AI engineer and when he was doing his internship at Tandem that he was really, really frustrated.
And he was very frustrated that he wasn’t doing any productive work, right? I think at that point, we were super eager interns, like, we wanted to learn as much as possible. But Cole basically spent the entire summer working on a project where he was doing a lot of prompt engineering for agents.
Specifically, like, his task was to help build out some of, like, the agents that Tandem was doing and some of the ML stuff that they were working on. And basically, what he ended up having to do is he had an eval set, and the eval set is basically a bunch of, like, test cases that an AI system has to pass or anything like that.
But all Cole’s, Cole’s job was, was basically taking that eval set And continuously just iterating on the prompt until it got to a point where it passed, like all the test cases for the evals and stuff like that. So then what that ended up looking like was he literally just sat in front of a computer and ins-instead of doing like the manual prompt engineering himself, he just ended up throwing the prompt over and over again into ChatGPT or any sort of LLM, right?
So he just threw the prompt into Claude until it passed, like all the eval cases. Okay, it’s like failing one eval cases or anything like that, threw it back into pro Claude, and then, okay, it’s like failing the system as well. And then eventually doing it over and over again was just super repetitive and manual, right?
I think that’s the point where it clicked in his head and afterwards, like when he told me as well. Because at a certain point, if he’s just doing this like automated loop work where he’s just, just throwing a prompt over and over again into Claude or anything like that we could probably build some sort of flywheel to be able to automate a lot of that process.
So part of his intern project and the first ever version of Lemma was actually built by Cole inside of Tandem, right? So he basically built an internal tool inside of Lemma inside of Tandem to automate that process. And that actually became the first version of Lemma.
Brian Thomas: That’s awesome. Found that gap, and love the story.
Cole obviously was just very frustrated, and I can see that. I can only just imagine him, tweaking a prompt and resubmitting it 100 times over just to get this through the testing. But that defeats the whole purpose of having agents doing automation if you’re doing a manual process with a human, right?
I did highlight that you guys really enjoy building things together as well, and I think that’s really cool. So thank you. Jerry, what sets Lemma apart is that you don’t stop at surfacing the problem. You trace each failure to its root cause and push a proposed fix back into the code base, even opening a pull request so detection and resolution happen in one workflow.
Why was closing that loop so important, and how do you earn an engineering team’s trust to let Lemma propose changes to their actual code?
Jerry Zhang: Yeah that’s actually a really great que-question, Brian, and it was one of the first things that our customer ever brought up in the feedback that they gave for Lemma and the stuff that we were loo– working on.
So just as sort of like a TLDR and like a high-level overview, Lemma is basically like the production monitoring platform for your AI agents. So we help you catch all the silent failures your observability or eval suite misses, specifically like the silent failures to do with agents being stuck in a loop or doing the wrong tool calls or forgetting or any type of memory issues.
Like the really tricky failures that usually teams have to manually wait for a customer to complain about it, or be the ones to kinda like dig through the traces themselves to figure out where the agents are failing, right? The ones where the agent executes successfully but produces the wrong output or some undesired effect or anything like that, right?
When we went to customers with like the first version of Lemma, one of the first things that we realized is that engineers’ trust and what they care about the most is actually, like, not the time it takes to notice these issues, but the time it takes to resolve these issues, right? With the first version of Lemma, we were basically surfacing a bunch of insights and issues to our customers, which was super valuable, but it wasn’t as valuable as where we needed to be because one of the really important things was that we just ended up giving more homework to our customers, right?
So a lot of our first customers actually hated Lemma because of the fact that we were surfacing a bunch of these issues, and at the end of the day, it was the engineering using the tool, right? So we’re basically giving them more homework or more tasks to complete before they can sign off or anything like that, right?
So it was actually frustrating to a lot of the engineers using the product that we weren’t closing the loop and kind of like solving the issue, right? I think that’s really, really important because not only do tools today, and especially with the capabilities of AI, not only does it have the capabilities to surface these types of issues, but it also has the capabilities to be able to fix those issues.
And I think that’s really important because tools today shouldn’t just be giving the engineering co– giving the engineers more homework to do, but rather giving them important homework to do, but also doing some of the homework for them, right? No engineer wants to spend like a Saturday night, like, trying to fix, like, agent errors and stuff like that and figuring stuff out.
If we can automate as much of the process as possible, then basically Lemma is the one coming to you, not only with the homework, but also with the homework complete. And the engineer is almost like the teacher reviewing the homework. And that’s like the analogy that we give and like what we do today to gain, like, an engineering team’s trust.
Brian Thomas: Amazing. Thank you so much. A- as you put Lem is essentially a production monitoring tool, right? Initially, that’s how it was developed catching silent errors, anomalies, et cetera. But what you said here to our audience was engineers really don’t care about finding the, you know, noticing the error, the anomaly, or, or catching it as much as they do the time it takes to resolve that issue.
And so the fact that you were working in the trenches with your initial customers, finding out you’re just giving them more homework to do was– I thought that was interesting. But, but I like now how you’ve refined your platform where it’s not just more of notifying, but it’s actually resolving the issue on the fly, which is really cool.
So thank you. And Jerry, the last question of the day, your tagline is deploy 1X, learn forever. A vision of turning static agents into self-improving systems that get better from real-world usage rather than relying on engineers to manually discover and patch every edge case. As you look five years out, what does a world of continuously self-improving agents actually look like?
And what has to be true technically in the terms of trust, right, for teams to let their agents learn and fix themselves at scale?
Jerry Zhang: Yeah. I think that’s a really, really great question as well. One of the first kind of like big differences and big value props of Lemma is that, and this is what I think is really cool as well, is Lemma is almost like the first ever version of a self-improving loop that we see for agents, right?
Obviously, the entire process isn’t automated. It’s not actually self-improving because at that point we would be at AGI probably. But we g-we get as close to a fully automated loop as possible where we have minimal human intervention, right? So on our platform, basically, Lemma alerts you about the issue.
As the engineer, all you have to do is, like, decide to triage the issue, which issues you want to be resolving, and Lemma kind of works alongside you to, like, fix those issues, right? So, like, Lemma pulls all the context about where the issue is happening and then passes it to the engineer’s coding agent to try to one-shot like a PR or anything like that.
And I think this leads into our larger vision of continual learning. I think every system in the future, and especially agent systems, should improve over time as it sees more real-world production data, right? I think it’s pretty intuitive that even as humans, like, the more we experience things, the more we’re able to learn and improve and grow from our experiences over time, right?
I think an agent is supposed to do the same, but what a lot of people don’t actually realize is agents actually degrade in performance, right? So the more the agent is used in production, the worse it gets because of things like prompt drift or model drift or anything like that. So we imagine like five to ten years out down the, the line, like we want to have like the first version of self-improving agents or continual learning, where agents themselves are able to identify where they’re failing automatically.
If they, they can figure out where it’s failing, then they can pull like the context to go and resolve the issue and then improve over time, and performance and accuracy gets better as more and more customers use these types of agents to gain trust with the consumer, right? The foundation of that, and I think where the first step is, like I’ve mentioned be- so far, is figuring out where the agent is doing well and doing bad in the first place, right?
So I think Lemma is kind of like working on, like the first step of walk, like working walking towards that vision, right? Because at the end of the day, the agent has to be able to tell, like what it’s doing well versus like what it’s doing bad in order to improve itself and be able to fix themselves at scale, right?
So we’re almost at first kind of like teaching a system or an agent how to be able to self-identify failures and figure out where it’s doing well versus where it’s doing bad. And then once we can kind of get that foundation set in, and that’s kind of like the bottleneck towards like all these manual improvements and stuff like that, Then you can basically sh-shift, like, from like a very human or developer-centric approach to a more AI native infrastructure where we see agents being able to do like continual learning and self-improving stuff.
I think like Lemma is almost building like the observability and the issue detection that becomes the foundation of that future vision.
Brian Thomas: That’s amazing. I love the vision especially y-you and Cole being so, young and, and, you know, not a, a ton of tenure, but you’re really just grabbing the bull by the horns here.
And just to highlight a couple of things, talk about what you’re doing with your platform. Obviously, right now, your value prop is it’s nearly a first version of a self-improving loop. But you know, over time, in the next few years, this will lead obviously into your longer-term vision where you’ll truly have that first version of an agentic self-improving loop.
Obviously keeping humans in the loop in that whole process, but I can’t wait to see the final product of this when, when things are truly running on their own, so to speak. So I appreciate your insights today. And Jerry-
Jerry Zhang: Yeah, super excited. Super excited to be building like the future of these types of intelligent systems and everything along those lines.
Brian Thomas: Absolutely. Love doing this stuff, and that’s why I get to interview people like you every single day in this technology- … which is awesome. And Jerry, it was such a pleasure having you on today, and I look forward to speaking with you real soon.
Jerry Zhang: Amazing. Thank you so much for your time, Brian.
Brian Thomas: Bye for now.
Jerry Zhang Podcast Transcript. Listen to the audio on the guest’s Podcast Page.











