Neural Frames founder Nicolai Klemke on building a profitable, bootstrapped AI music video startup to $5M ARR with 12 people in Berlin
Aug 10, 2026 · Full transcript · This transcript is auto-generated and may contain errors.
Featuring Nicolai Klemke
Speaker 1: So fun. Let's bring him in. We have Nicolai from Neural Friends, the founder and CEO. Let's switch over.
Speaker 3: What's going on?
Speaker 1: How are doing?
Speaker 2: What's happening? Welcome to show.
Speaker 3: Hi, guys.
Speaker 1: Great to meet you.
Speaker 3: Great to meet you. Very exciting to be here.
Speaker 1: Thanks so much for hopping on. Introduce yourself. Tell us a little bit about the company, How long you've been building, the size of the team, where you're based on.
Speaker 2: Well, yeah. And a little bit more backstory. I think the first your first attempts at making music videos, I think you were doing manually? Maybe. Yeah. But then but then at some point it switched over and you were just flipping them over to me like every twenty minutes and I'm like, how is how is how is John doing this? Yeah. Now we know. Yeah. He's using your product.
Speaker 1: Yeah.
Speaker 3: Now we know. Now we know. John, first of all, props for for the music video. Amazing what you got platform and the and the lawyer. The lawyer.
Speaker 1: Real real career ahead of me as music AI music video director. Sure. I think
Speaker 3: so. I think so.
Speaker 1: Well, no.
Speaker 4: I mean,
Speaker 1: we're going to this about, like like, it it it's not it's not a broadly, like amazing video. It's only funny in this context of like there's like 50 people that think that's funny that understand Quinn Emanuel, but you would never be able to underwrite making a mute full music video for a random corporate lawyer. But you can now, and I think that's what's actually interesting about this product and AI video generally. But but give us the history of the company.
Speaker 3: Sure. So, yeah, I'm I'm Nicolai, founder and CEO of Neurofrints. This started three and a half years ago, first as a kind of indie hacking solopreneur project. I just kind of played around with Stay With The Fusion. So I'm I'm a I I have a long time history as a musician, as a passionate hobby musician. I have a PhD in physics. And then I kind of wanted to bid something myself in the real world and not only in in physics labs. Basically, I was looking for a startup to bid and in the end of twenty twenty two played around with AI animation based on stable diffusion, which has just kind of come up. And I deeply fell in love with creating these animations. It was it was so much fun. I had a model fine tuned on myself, and then I, you know, it was an image to image to image loop, and I could just steer the animations with my own words. And that was so much fun. I loved it. I thought, like, I cannot be the only one who who enjoys this. I and I didn't really have a clear use case in mind, let's say, but just hacked together a prototype launched in January 2023. It kind of went viral on Hacker News and and read it and then just grew from there. Fast forward to today, we are a platform focused on music videos. We are a team of 12 people at the moment in Berlin. Yeah. We are bootstrapped still. We are profitable. I I don't know if you hit the gong for
Speaker 2: bootstrapping definitely hit the gong for profitability. Those are our favorite gongs to hit.
Speaker 1: Yeah. That's the best. Yeah. This
Speaker 3: is my my best dream coming too, John. Thank you very much. I mean, we're doing 5,000,000 and your run rate. We grew six on year and so we are we are very excited for what's
Speaker 1: to come. Yeah. Yeah. With 10 people, I'm sure it's gone very well. Yeah. So I mean In the castle. In the castle. So, yes. I mean, so many questions. First, I'm I'm interested in the positioning of the product because I I found neural frames by actually going to chat GPT and saying like, I have a Suno song. I want to make an AI music video and I want to do not a lot of work. Like, what do you recommend? Recommended neural frames and then I went there. Good. But I'm interested in in how how opinionated you think you need to be in the product, how narrow you want to stay focused because mid journey is a great example of a company where it feels like it has David Hole's fingerprints on every image that's generated and that could be a criticism because it's not this broad platform that's AGI doing everything, but there's an immense amount of value in having an opinionated product at this moment in time. And I'm wondering, do you think you'll stay there or do you want to broaden out, or or do you want to go more opinionated?
Speaker 3: Great. That's a very interesting question. I mean, so we are very focused on the music video use case, which is already kind of a niche by itself. Right? We are very deliberate there because we are bootstrapped, there's no purpose in in competing with, like, generic general purpose AI video companies out there. Right? Within the AI music video realm, it is definitely our ambition that you should be able to create kind of anything. But, of course, we have kind of our harness. I think you called it harness before, and this is also how how we call it internally. We kind of build our harness so that people get kind of high quality, at at least like the highest possible quality with with the lowest amount of clicks out there. Yeah. I They're kind of trying know to follow what they what they type in.
Speaker 1: You could actually use the service for non music videos. I I found a very weird story in my hometown of like a fraud case that there was only like one article about this this particular case in like the LA Times. But I was able to take that into ChatGPT, turn it into a deep research report, turn that into a script for a Netflix style true crime documentary, then take that into Suno, make voice over and then I took that into Neuralframes and it basically made like a AI generated six minute Netflix documentary. Neural frames originally initially got a little confused because it was like expecting beats, I think. And like Right. It was having trouble cutting it up. But it wound up working just like once I refreshed. And so there there's a world where this like vertical short form storytelling could get there. But it does feel like there's still a really like the the magic is definitely in if you're not going for the funny shock value of like the content like, oh, it's really polished about something that does not deserve a music video, the quality bar is a lot higher. And so I'm interested in how you're assessing progress across the various video models. I did try CDance 2.5. It's good. I don't know if it was my prompt, but it was still very clockable as AI imagery. And so I'm wondering about how you're thinking about the the progress in AI video models, the cost, and what you're seeing across the landscape.
Speaker 3: So, I mean, first of all, one year ago, we were, I think, to the best of my knowledge, the first music video tool out there that had character consistency.
Speaker 1: Oh, yeah.
Speaker 3: Right? Mhmm. Just to to to tell you where we're coming from. So, like, one year ago, it was impossible to have character consistency, basically. Then we added three character consistency, and now, I mean, the lawyer was with all the criticism that you had, it was pretty consistent person, and I and we're actually very proud of, like, the the character consistency. You can even upload yourself and and create
Speaker 1: Yeah. Yeah. And it's easy. You just take a couple screenshots, upload it, it creates a character, and then and then you just tag it, and it's very, very simple to actually, like, create some consistent character.
Speaker 3: Of course, like, if you imagine, I mean, AI video models create usually five, ten second clips. Right? Now, c dense 2.5 can even go up to thirty seconds, but that's very new.
Speaker 1: Okay.
Speaker 3: And our kind of challenge is the the main challenge of music video generation, okay, you have a three, four, five, six minute song, and and we need to create all the images and videos that kind of match the beat, match the mood of the song. Right? Mhmm. Match the character, tell tell a story. Mhmm. And and so it is actually surprising that the AI video models themselves are important. We can do cool stuff with good new models, but it is surprisingly, I would say, not the most important aspect of this whole thing. The the whole harness and image generation is is at least equally as important. And, of course, we're continuously working on that and making, I think, also very good progress. I I think this this video that you showed would have looked very different three years three months ago already.
Speaker 1: Oh, yeah. Yeah. Yeah. Totally. How do you think about the trade off between, like, UI and and sort of really surfacing the feature versus baking everything down into a text box. Because there's a world where I could have just gone and said make a music video about John Quinn of Quinn Emanuel. I did have a few steps where I I went to Suno, generated the music, gave that to you guys and then uploaded the photos. And you could sort of do all of that in one huge box and huge prompt, but obviously you get higher quality as you layer in more user interaction. And you're just surfacing more of those. Like, I didn't know the character consistency was an option until there was like a little button there and I was like, oh, what's that? Then I see and then I go down that flow. If it's just a blank box, it's harder to do the user education.
Speaker 3: Yeah. This is a very interesting question and and one that we are debating at least every few days here in in the team. This is also changing a lot. I still think there is a lot of value in in opinionated UX somehow to to guide the user along certain steps. Right? Mhmm. And also to allow great editing in the end. Yeah. So we we I mean, our power users, they spend a lot of time on the platform refining each scene. Right? We have a timeline where you can drag things around. We you can change the characters. You can change the outfit and stuff like that. I just personally feel in a in a chat based product that would be a bit more difficult to do. But let's see. I mean, we might experiment with that in the future.
Speaker 1: Okay. Give me a lightning round of the video models. I want your reviewexplanation of what they're good at or what they're not. Let's start with C Dance 2.5. It's now on neural frames. What is it good at? Is it too expensive? How should I think about that model?
Speaker 3: It's an amazing model. It's probably the best model that exists at the moment and can create up to thirty second cohesive clips, and it has, a multi shot feature, so it it doesn't create one long thirty second clip, but multiple shots at the same time, which NeoFrames kind of does for our models already. But but in in CDense, it's kind of integrated. You can also upload lots of references, so lots of objects can reappear across scenes. All of this is very new. Yeah. In theory, there's also audio input in in this model, which sometimes works better than Yeah. And sometimes doesn't.
Speaker 1: What about code?
Speaker 3: It is also it is also probably the most expensive model that exists, and and cost is cost is definitely a huge issue. I mean, video models are are quite costly. Right? And Yeah. Cling, it's amazing as well. Cling three is a great all rounder kind of also, it's it's not the cheapest, but significantly cheaper than Z lens two and Z lens 2.5. Yeah.
Speaker 1: What's going on with the VIO family of models from Google, Gemini? Very impressive at various moments, but I haven't seen them really be harnessed in the same way as Cdance recently.
Speaker 3: I think the latest releases of of Google were Nano Banana Light, which is a very good lightweight image model, actually, which is also very useful for for several steps in in our workflow. Okay. Not the highest quality ever, but kind of affordable and very fast. Sure. And then they also released Gemini OmniFlash, I think it's called. It's a it's a video model that allows also editing. It's also very good, but in our experiments, doesn't quite reach the quality of CNNs too.
Speaker 2: Mhmm. Have have a lot of marketers discovered neural frames? Can I imagine people just saying like, yeah, well, made it for music videos, but, you know, if you wanna make a Facebook ad, you know, and you have you want think
Speaker 1: Facebook ads should be music videos, more likely to stop scrolling?
Speaker 3: Right. We had an we had a cosmetic brand using Neuroframes. They they kind of created an AI influencer that then created songs with one of their AI music tools and and then music videos with Suno also with with Neuroframes. We're also seeing musicians themselves running ads on on Instagram or Facebook that then link to their Spotify playlist. Mhmm. So that's also an interesting use case, I think. Yeah.
Speaker 1: Yes. That's fun. Well, thank you so much for coming out.
Speaker 2: Yeah. Very cool.
Speaker 1: Do have anything else?
Speaker 2: No. Have a great day. For Thank providing so much entertainment.
Speaker 1: We're having a good Definitely time with
Speaker 2: definitely tens of thousands of dollars worth of entertainment. I don't know how much how much you're you're you're charging us, but it's been amazing.
Speaker 1: It's not cheap, but it's very affordable for what you get. Like, you know, these things are still
Speaker 2: Yeah. Belly laughs aren't Priceless. Yeah. There's no such thing as a free belly laugh.
Speaker 1: That's true. That's true. That's true.
Speaker 2: Great to great to yeah.
Speaker 1: And congratulations on the progress. Keep it up and come back on the show when there's no more there's more
Speaker 3: Big fan of you guys. Thank you so much for
Speaker 1: you soon.
Speaker 2: You're the man, dude.
Speaker 1: Have a good one. Let me tell you about Cisco. Critical infrastructure for the AI era. Unlock seamless real time experiences and new value with Cisco. Up next, we have two
Speaker 2: I'm surprised that Mikey hasn't come in with the maxed out offer yet.
Speaker 1: Yeah. Yeah.