Sunday Robotics achieves 99.1% success folding laundry across 785 zero-shot attempts, introduces 'solve' framework for robotics benchmarks
Jul 20, 2026 · Full transcript · This transcript is auto-generated and may contain errors.
Featuring Tony Zhao
Speaker 8: We're actually moving pretty soon. But yeah, don't really move much from this spot.
Speaker 2: Well, we appreciate you hopping on the show.
Speaker 1: Awesome. Yeah. Great to get the update.
Speaker 8: Thank you so Have
Speaker 5: a great rest
Speaker 2: of your day. We'll talk to you soon. Up next, we have Tony Zhao from Sunday Robotics. He's back on the show with an amazing update with some remarkable benchmark. 99 There
Speaker 1: they are.
Speaker 2: Zero shot success, folding laundry across 785 autonomous attempts. They're calling it laundry super intelligence. How's it going? Is it solved? Good. Is it is laundry solved?
Speaker 9: Yeah. We call it LLM.
Speaker 2: Oh, yes. Which
Speaker 9: means large lounge laundry models.
Speaker 2: Perfect. Perfect.
Speaker 1: Move the goalpost. We gotta move them.
Speaker 2: No. Give us the broader update. Like, what what is act two and what wait. Like, is this a discrete training run? I mean, are you into actually commercializing this? How much is in the lab? Like, give us the broader update, then I'm sure there's a bunch of ways we can go deeper.
Speaker 9: Yeah. I think this is actually a really packed update with many things that we want to share. I think the first thing is this is the first time we can train a policy, a model to both generalizable and reliable. What it means is that by generalizable is that what you see is what you get. The performance you are seeing will be the same if we deploy the robot into your home because these have been tested in places that are unseen that the model generalizes into. And what we mean by reliable is that, like, after filling, like, close to a thousand garments, the success rate is, like, ninety nine point one percent, which is very, very reliable and is the most reliable one to date, and it's handling, like, very diverse amount of garments. So this is kind of the more the most literal part of the update, but I think there are also many, like, new ideas that we introduced. One of them is that it turns out that when we scale up pretraining, you can do one shot learning with these models. And what we mean by that is that you can teach the robot to fold your shirt in a new way with one demonstration.
Speaker 2: Mhmm.
Speaker 9: And the model can extrapolate that to an unseen shirt on an unseen bed. It actually, like, understands it. So this is something that is, like, really, really surprising that just emerges from, like, scaling up the training the pretraining, both the data and the compute. The other thing that we are, like, really deep into is to introduce the idea of a solve. So I think this is there has been, like, a huge amount of confusion around, like, what is 99% this, 99% that? But 99% can mean very different things when the scope and the adaptation budget is different. Let's say if you're doing 99% folding one shirt in one room, this is very different from being able to fold any shirt in any new places it deployed into. So what we're trying to distinguish here is that the performance, let's say success rate, only matters if you very well like, very nicely define the scope, which is like what is a long tail of things you want to handle, and the adaptation budget, which is another way to say, like, are am I allowed to train on my test set? Am I allowed to train on the deployment situation that I'm going to do? And only after these boundaries are established does a success rate make sense. And one example of that is, like, if you think about self driving, doing 99% in a closed course versus on highway versus on city streets, they're, like, very, very different challenges. And so I think we as a field in robotics are at a point that we want to clarify these to, like, better be able to track the progress in the field.
Speaker 2: What does prompt engineering look like as this gets deployed? Like, can imagine, you know, you go to ChatGPT and you say, like, write me blog post and it'll be like very mid and then you give it an example of how you write and it gets a lot better. And there's a lot of style transfer examples. You feed in an image to mid journey and say, I'm looking for generally like this but use it, make it a dog. And I'm wondering if there's a pattern where at least in interim, in the midterm, you or in the medium term, you will say, Okay, you know, practice is when Memo comes to my house. Give it a couple examples of how you like your socks folded, you know, show it the long socks and then also show it the short socks to give it a couple reps and then you'll be good. Is there sort of even though obviously you're doing a lot to make the product just work out of the box, is there going to be some sort of like last minute ad hoc on the fly fine tuning that's essentially happening in the home?
Speaker 1: Well, yeah. A good example is like folding all the clothes is one challenge and then putting them in the right places is another Sure. Which, yeah. Everyone's had the experience of being like, okay. Well, not not everyone. But like, I'm not home. I'm not home when my when my laundry gets folded and sometimes it gets put somewhere I don't know where it is. I want it to be put in the right place
Speaker 2: Of course.
Speaker 1: The right time.
Speaker 2: Yeah. Of course.
Speaker 9: Yeah. I think what you guys are highlighting is really the the need of personalization
Speaker 2: Mhmm.
Speaker 9: For homes. And that's actually a hallmark of intelligence. Right? Is that if you're so smart, you should be able to learn from one single example Sure. Which is what we thought for, like, the front engineering side.
Speaker 2: Yeah.
Speaker 9: I think this is something not something we're going to roll out this year, but maybe, like, either super late this year or, like, early next year
Speaker 5: Yeah.
Speaker 9: Of, how can we allow users to customize the behavior of the robot? Let's say, I just want my laundry to be done this way or I want it to be placed into the cabinet in that way. Want my home to be organized in this way of, like, a specific way of laying out, the living room. And can you do that just by one video, one photo, or, like, one demonstration? This is actually what this research update, the one example, one shot learning is about. That it turns out as we scale pre training, that capability just emerges. And, of course, there are work that needs to be done there to make it a actual shippable feature, but it's really, really promising to see that it's even possible.
Speaker 2: How do you think if you're really successful, if there's a fast takeoff of of this product category broadly, hopefully you with you at the forefront, how do you think it will change the design of American Homes? Like, can imagine some people have laundry rooms on the 1st Floor, bedrooms on the 2nd Floor. You don't have the ability to climb stairs. I could imagine people saying, hey, maybe I'll remodel and put a laundry room upstairs. But have you thought about the knock on effects of like if somebody really leans in, you know, they're getting the electric charger for their car, they're modifying their home for the best experience in the future. How will things change?
Speaker 9: Yeah. I think there are a couple of things that are really interesting here that you mentioned. One is that like, I think we're definitely in a period of time that the capability will take off.
Speaker 2: Yeah.
Speaker 1: In a
Speaker 9: sense that like, when we work on this laundry task, we didn't make any assumptions about the laundry. The recipe itself is extremely general. Mhmm. So there's really no fundamental reasons why we cannot work on a thousand of these skills in parallel. And which means that, you know, the capability is only exponentially growing scale as if we can scale data and compute accordingly.
Speaker 2: Mhmm.
Speaker 9: And I think, as you said, like, the the the revision of this is that if the robot can do so many things in your home, a lot of them are, like, maybe across different floors and across different places. Is there ways to modify the environment? Is there ways to modify the robot? And I think this is also where our technology comes in, which is the way we design a data collection device is agnostic to the hardware, That the same dataset, the same model can be used to control memo as it is right now, but maybe it can control a legged version of memo in the future.
Speaker 2: Sure.
Speaker 9: But we're actually super open minded about form factors, but within the ecosystem that we're going to build, which is going to be beautiful high quality hardware.
Speaker 2: Got it. You said there's a thousand tasks you could think of. I can think of like four things that a robot could do in the house. It's like laundry, cooking, dishes, cleaning generally, maybe like gardening. Like there's Locking got to
Speaker 1: be the lock doors at
Speaker 2: locking the doors at night? Yeah. I guess there are some more. But how how long is your task list? Obviously laundry really stands out as like something that very few people enjoy. It's very obvious and it seems very unbounded. So if you can solve that, you can imagine it also being like I would trust that robot right now to throw a lock. That seems easier than than folding a shirt. But how how long is the list of of tasks that you want to that you want to knock down over the next couple years?
Speaker 9: I think there are almost like two parts of this question. Yeah. One part is that what is the minimum number of tasks that a robot needs to do to justify its own existence Mhmm. That people love having it in their homes. Yeah. And as you said, I actually think the list is pretty short because the annoying chores are just that many. Right? Yeah. They're pretty repetitive. You just need to learn it, and, you know, you just keep doing that over and over again, like folding shirts or like loading dishwashers.
Speaker 2: Mhmm.
Speaker 9: But I think I think that's one part of the equation. But the other part, which is what we noticed, is that I think we're also having this longer term goal of can we get to a general intelligence? Can we solve the, quote, unquote, physical AGI?
Speaker 2: Yeah.
Speaker 9: And what that entails is a way longer list of tasks. And what we are seeing just like, I think home and this objective, actually aligns quite a lot. We call it, like, research market fit. That's the research that goes into the product also advances on the general intelligence side.
Speaker 2: Mhmm.
Speaker 9: So so overall, I think the there are, like, obviously infinite amount of manipulation skills. Even for one task, you can do it in many different ways. And I think we'll first cross the boundary of making a product exist, but we'll keep going towards the north star of being able to solve any task with, like, very, very few demonstrations.
Speaker 2: Mhmm. Very cool.
Speaker 1: I can believe the scenario where a winning robotics company starts off by just doing something cute like folding laundry and and then, you know, ten years later, it's doing everything.
Speaker 2: Yeah. No. It's a good place to start. I love it. Well, congratulations.
Speaker 1: What's your what's your what's your your cooking timelines?
Speaker 9: I think cooking is such a fantastic task. Like, everyone can have Gordon Ramsey level of cooking in their homes. That would be, like, crazy. Like, people pay, like, so much money for it. But at the same time, I think laundry is also one of the tasks that are, like, difficult. Right? Like, if you, like, accidentally mess up, you need to clean up after yourself, which is okay. So I think we think about laundry as this almost like a like a final boss that the upside is so high, but
Speaker 1: Where you think laundry is harder than
Speaker 2: the opposite because
Speaker 9: Oh, sorry.
Speaker 2: Can't shatter. Yeah. Cooking feels like the final boss because you can break glasses. There's heat involved.
Speaker 1: Yeah. There's like Yeah. Come back into your kitchen. It looks like
Speaker 2: Whereas like, if I have a shirt and the robot messes up and and gets the shirt all bundled up, it's just a messy shirt. It was already a messy shirt.
Speaker 1: Not if it's chrome hearts.
Speaker 2: Yeah, guess. But if it but but if you come back and it's like, okay, you actually like broke a glass of wine and you like you broke up a glass of a bottle
Speaker 1: of olive on the floor.
Speaker 2: Yeah. There it's a little bit riskier. So but it it seems like it will transfer pretty well, at least in