YouSaid · the spoken record

Jasper Paterson

lines on the record
58
first
2026-06-15
most recent
2026-06-15
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. I think for enterprise, again, we have some sales forms they can fill and then come talk to us because we see that there are many differences in what different companies want. Some companies care more about editing. Some companies want more like on the marketing side to automate some of their ads. And we should talk first and then figure out what's the best solution for them.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  2. There is actually a model top if you log into Idogram, there's a model tab and then you can go and upload your images and train your model. It's a little more expensive, it's 60 bucks for two model training per month, but they think for professionals, that's totally worth it.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Companies who want more control and data privacy and sovereignty. So we would love to work with other enterprises as well across the stack. And then the best way would be like we have this email partnerships at Idogram. You can DM me on Twitter or LinkedIn and I'm very active on both platforms.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  4. First of all, we would love to work with more engineers, like cracked engineers. We have a very tiny team. You see what we were able to produce. It's such a tiny team. And if you want high agency, if you want your work to matter and part of and you want to be part of the academic and open source ecosystem, then this is the perfect time to join us. Now, in addition, enterprises see the potential and we would like to work with the most creative brands out there to help them produce the best designs, produce most provocative ads. And also we would like to partner with other startups or other companies at different levels of the stack. This is open weight so we can make it win-win and we would love to offer a different option to come.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  5. Love to hear if there's anyone who's interested in working in Ideagram or working with you guys as a customer or to fine tune a model what's the best way for them to get in touch with you and the team.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  6. It's actually kind of alluding to the editable model that I'm talking about, and we've had a lot of back and forth, or we're going to have our own JSON for different text elements and buttons and stuff, or are we going to use HTML? And it seems like HTML makes more sense just because these large language models have already been trained on HTML as opposed to us introducing a new JSON structure. But I would say to answer your question, that representation needs to be easy for the language model with the particular design that we have right now, which is a language model, does some expansion of the ideas, and then the image model takes those expanded descriptions and turns them into images.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  7. That's a very good question. So, in general, the recipe for building more powerful models, in my opinion, is making the task as a straightforward as possible for the diffusion model that is specify the exact details of the image. And so now if you kind of make that extreme, then it becomes the pixels themselves. So the fusion model doesn't have to do anything. Now, the catch is what we would like to do is to get the language model to produce that intermediate representation. And language models, as of now, they aren't very good with continuous output. They aren't very good with kind of pixel values or very high dimensional vector representations. So I guess the constraint here is that the representation has to be tokens of like, I mean, depending on your language moral power, maybe

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  8. Yeah. So I guess the question also becomes like I kind of asked you the representation question before, which is here's a JSON representation. As an artist, obviously, if you abstract it out, you know, far enough, all the lines are pixels. So you could say that the composability is on the pixel level, which is actually different from the diffusion representation. It's like denoising and, you know, they operate in a different space. Where does this lead to if you travel down the JSON path granularly enough? Does it lead to pixels? Does it lead to SVGs? Does it lead to like language or something else?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Yeah, so we try to be. Enabling very different styles, and that was one of our goals. We still want to produce tasteful output, but that doesn't mean we have to force a complex output to you. If you want a minimalist design, you should be able to get that. Actually, our minimalist is too minimalist, in my opinion I was saying we should ban devote minimalists from the output, but you get what I mean. It's like the model can do many different things and that's by design We know

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  10. I think when you talk both about the design and the art, it really brings back the taste point you made earlier because for so many sorts of designs, you're trying to communicate some sort of idea, whether it's like an infographic or an ad or a logo or whatever. And it kind of needs to stand out and be distinct to you or have a unique style. And I've totally noticed what you said. A lot of the frontier, the historical frontier image models, like you're scrolling the feed and you just see like, I've now seen this style like 50 times. I've seen it a hundred times. It doesn't catch my eye anymore. And it feels like now when I prompt ideogram four, I often get something that like makes me stop and be like, wow, this is different than anything I've seen before coming out of an image model. And like, this is doing an amazing job of both communicating what I want to communicate and also holding someone's attention.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  11. The score very highly in the leaderboards. They don't have a lot of kind of design variation. They always produce the same exact look. And I believe that's because they did a lot of reinforcement learning training. I actually have done very little reinforcement learning. So this is a very raw model. Now with that, the outcome is you need to be much more precise with your prompting, but you can get a lot of different styles from the model and people seem to be very excited about that aspect of the model, especially in the art community.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  12. In terms of the things we've seen from the model, again, design is something that's coming up a lot. And I saw a tweet before coming here that somebody said, I have no design training. And I got this design in two minutes. And it actually looked really, really nice. That was one example. And then people are really excited about the art possibilities because this model was trained with very unique style description as part of training. We actually stripped that from the JSON prompt because it became too much. But the model has a lot of artistic possibilities and like many different styles are embedded into a model. And if you've seen some of the frontier models, actually,

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  13. You need a UI. You need a UX to be able to go and edit, and whether it's regional editing or text-based editing, I think at the end of the day, you want your Canvas, you want to be able to point to things, and then you want to also talk to it with national language. It's actually very hard work because models are changing and now you're also designing the user interface at the same time. Kudos to the best designers who understand how these models work and are trying to figure this part out. There's still a lot of work to do.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  14. What's interesting is, yes, we have these agents, a lot of them live in a chat bot, and that's not enough for iteration, unfortunately. So I think what's unique about that, you can really scale creativity, like kind of give it some high-level direction and ask it to go and explore many different approaches and come back with hundreds of thousands of designs that can be easily looked at and then you get a better sense of, okay, I want to explore more in this direction. So the language model interaction allows for kind of very large scale exploration of creative possibilities. But once you know what you want,

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  15. I guess for the API business, it's so much of a design iteration. It's a long tail of a design. It's no longer just you prompt something, you get an image and color a day, right? It's so much of a get an image, use the edit model to edit it, see if it works well. If it doesn't, get another image with the JSON flop, which is easier for control. What are the net new use cases you have seen after launching a model of how people do compose different API calls on the ideology gram model?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Is when you want to release a new feature, you can go into your agent. Ask it to connect to the API and generate a bunch of images, and then you can go and find the best ones. And like in a couple hours, you have your landing page up and running. So we are very, very excited about agenting workflows. We think there's just at the beginning, to your point, we need evaluation as part of the loop. We don't want to be to have to look at every image. And then editing will be part of the agenting interaction too. How we exactly want to compose these different pieces to accomplish a goal is still to be discussed, but we have API and we have MCP and we really want to enable the agentic workflow. And we think every company is trying to figure it out as well. So we're really excited about that.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  17. And then say, so there is a lot more diversity in the visual world, and that's very exciting for customization. And there are a lot of unique ways of interacting with the models that kind of goes back to your earlier question, a lot of 3D manipulation. So for that kind of use case, the input will be some 3D representation of the joints or position of the objects. Then you may have a completely different stylistic variation with the style being the input. So there are a lot of different types of interactions that you want to enable. And for that reason, it's much different from the language space where the input is always text and like you can kind of more or less convert everything to text. We're so excited about agents. We have our own MCP. We use them a lot internally. What's really exciting.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  18. And your earlier point, I want to say something which is we seem to compare image generation with language models a lot. And for example, in the language model space, even though customization exists, it's not that every company customizes their language models. But I think that actually misses the point. When you look at Visual representation of a brand, you immediately recognize the differences between brands. But if you look at the written communication, can you say, oh, is this Andreessen Horvis or this is Sequoia? I mean, you probably can tell.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  19. With the JSON prompting and editing and model fine-tuning, the compose possibility aspect of the model is just huge. There are so many ways you could customize it. One hot topic in the industry, in the research community is a Gentex loop for creative tools, right? So it used to be the creativity tools. The consumption layer is always a UI. As a human, I look at it and then I make modification. Now, so much of that may become like an API request, like the agent makes. How do you see that? What would the API entail compared to how humans use it?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  20. But then customization gives you really freedom to not prompt at all, right? Because sometimes it's very hard for you to say, oh, I want to get inspired by this single image and edit that one, you may have some general style in mind and want to ideate in the context of that general style, or you may have a character that has many detailed degrees of freedom or characteristics, the left side, maybe different from the right side and like certain outfit. And it's very hard to really put all of those images as the input to your editing model and it often fails. So we think customization can give you a lot more powerful adherence to your characters and allows for an easier iteration on ideation. So I agree with you. I don't think Are mutually exclusive, but they're both very powerful.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  21. Yeah, I agree. I mean, I think editing is very powerful. We both agree, we all agree when editing launched last year, so many new possibilities opened up. And the nice thing about editing is it's quick. You don't have to train a model. You just take a style or an existing image and you make some changes. And it's part of your iterative workflow because with every creative that we worked with, it's never one-shot prompting. It's often, okay, you get something and then you're going to go and fix certain details after the first generation. And that's for editing.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  22. One of the things I think people have been talking about a lot, both for enterprise and actually for consumer, is fine-tuning versus image editing. I actually think they don't necessarily have to be competitive. Like some people use image editing as a way of fine-tuning. Like they say, take this image and put it into the style. Others think it's much more efficient and consistent to just fine-tune a model to generate in that style. I know you alluded to wanting to release an editing model further down the line, so would love to hear how you guys are thinking about that.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Who have certain names. And so our annotation team gets involved and spends a lot of time curating and printing data. And we think depending on your size and your budget, you should still be able to customize the model, maybe use the open source. At the low budget, and then you can come and talk to us so that we can build a model for you at the high budget, but then depends on really the ROI that you have in mind for your models.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  24. So, one thing I should say is We will work with the open source community to make it as customizable as possible. But because we are the model developer, we have some secret sauces that can make it even better. So what will happen is there will be different ways of customizing. One is in the open source, based on the quantized model that's already released. The other is we already have a product that allows you to customize by just uploading certain number of images to our custom model training app. And we haven't released the 4.0 version of that yet, but we are hoping to release that as well. And then they kind of third category is when enterprises work with us and we really describe these detailed prompts for every image. We work with their design team to understand what wars they want to use because each team has different set of keywords. Each company may have certain mascots.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Yeah, yeah. So some of the enterprises we work with are very sensitive. They don't want to talk about using AI in their visual side of things. But what we've seen over and over again is Companies come to us and say we tried these generic models and they don't meet our design bar. They don't follow our style. They don't follow our brand guideline. And once we train custom miles for them, they are like, wow, this understand my brand DNA now. We can use this for design ideation or we can use this for marketing. We thought with open rates release we can give a glimpse of customization to you know developers within enterprise and kind of scale this side of our business further and we're really excited about that.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Were you getting a lot of demand from enterprises when it was a closed model that wanted to fine tune it on their own data? Were there use cases where you're like, we really need to open source it now because there's so many cool things that people want to do?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Style, the texture of their canvas, and really get 2K output, and hopefully make that part of their workflow and, you know, be augmented with AI and be a lot more productive and creative. And we actually have worked with some artists in residence who said to us, okay, this at least made me 3x faster in making this comic book. So that's one frontier. Another frontier is enterprise again because it's not about how good a model is in the general sense. It's about how good is this model for my use case, right? At the end of the day, I may not care about many of the general purpose use cases. I may not care about character consistency, for example, as an enterprise. But even though the small is small, they think it can be the best model for particular use cases. Whether that's a certain artist or whether that's an enterprise.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Terms of the research team, it's an interesting question whether you can focus on a very small, narrow field in image generation. I sort of believe that you need a general understanding of the world in order to even be good at logo generation or be good at illustration style. But once you have a general base, then you can customize the model for certain use cases and it can be the best at that particular use case. So we're really excited about customization. We think that's a new frontier. And again, that's an important reason for releasing the model with open weights. Every artist who has at least like 50 pieces of art or hopefully like a little more, they can really customize this model to the new answers of their

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Right, so one thing that you kind of alluded to is the fact that this can run on consumer GPU now. And we think there is a new frontier that you do a lot of editing on your phone, a lot of image generation on your phone. And it's not only about pushing quality at 100 billion parameter, 1 trillion parameter range. We think it's really important to have small models that can run on device. Obviously a lot of companies care about privacy. And we are really excited to partner with the industry to push the kind of small model size quality further.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  30. I imagine because it's a smaller model, as you mentioned, it's harder to win on scaling and counting number of trips, but it is possible to win on a specific domain or optimizing for a different thing in a different domain. So what was the trade-off for the research team when training this model to decide what to focus on?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  31. Model primarily because we think there's just so much to do. We think now is actually a good time for us to scale given the quality of the model at 9.3 billion parameters. You should imagine what if this model is 100x bigger under a mixture of experts architectures that don't make the model necessarily store, but they make the model a lot more powerful. So I think that's one new frontier for us to kind of scale this model 10x, 100x.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  32. We focused on the details of the model and we know we can win on scaling. I don't think I used to work for Google. I don't think even if we raise 10x the amount we've raised so far, we can beat Google in terms of the number of chips that we can dedicate to each model training. So instead we focused on innovation. We think there's so much more to do to innovate. We are also focusing on differentiation. I don't think a lot of labs are focusing on design. In graphic design in particular, editable text that I'm talking about. And then we also decided to go open weight to really partner with a lot of other platforms to be at least another option for people who care about design. And so yeah, so we focus on the small

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  33. And then I guess one thing we were always wondering is that this release, the open source model, is so small. It's 9.3 billion parameters, like previously the SOTA is probably like 80 billion parameters. So it's like nigh x of a difference, and then you can run it on a single GPU instead of having a lot of compute footprint. Which really opened up opportunity for people to use it. So the question is how did you do it?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  34. Yeah, we just can ultimately work with all the arenas. We hope all of the arenas will improve in detecting the nuances of images and image quality. But we care about our own internal evaluation and unfortunately we see that AI is not very good at doing the actual taste evaluation yet so we work with designers and we have side by side comparisons between different versions of the model as well as other models to really push on the taste. So we really care about taste. I think there's still so much more to do obviously.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Right, so we care about a couple things. One is graphic design in general again, text engineering is part of that. We think Basic graphic design is everywhere. Like we go, you go in a city, you open your eyes, you see billboards, you see storefronts, they all have text. And actually it's much more important than, I guess photography is part of graphic design, but like graphic design is actually the frontier for a lot of business use cases, for storytelling. So we definitely focused a lot on graphic design since the release of our first model, which was good at text rendering. And in addition, we think taste is extremely important. And we really want our models to have taste. And it's very hard to explain it. What exactly taste is? One element of taste is kind of going outside of the norm a little bit and not conforming to the average opinion, which is a little against being on top of the leaderboard.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  36. Because the JSON prompt describes every detail in the scene, you can take it and change one element in the scene, and that results in very, very consistent output. So you could be describing like a tiny detail in the corner of the image and then leave everything else the same. And we think this also has a big implication on editing. We haven't released our editing models yet, but they will also utilize the same JSON prompting approach. And it's just more control. And with layout as well, you can imagine for every brand, you have brand guidelines in terms of, okay, the size of text, the font of text, and we think this kind of foundation allows us to really get into a lot of the enterprise use cases.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Maybe zooming out a bit, the world is changing, right? I'm sure your work has changed a lot. My work has changed a lot. I'm actually writing PRs now, which is very exciting. Or in collaboration with AI, AI does the writing. And then for creative professionals as well, the world is changing. They are very excited to. A large number of them are very excited. I think more and more of creatives are excited about adopting AI and they see the potential. They think ideation is still the most important part of the creative process and humans are very good at putting context into these models and their understanding of situation creativity will result in the best ideas. I'm very excited about the future where every kid will have these models at their fingertips and then they can be much more creative and we're going to experience a much more beautiful world. Now as it comes to JSON prompting again

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  38. Exactly. And then I think everybody else does it too. OpenAI does it, Google does it, but then they don't give you the actual input to the model. But again, for professional use cases, you don't want to just roll the dice and then get some other completely different image interpretation of your prompt. We show you the actual input to the model as well in the JSON format. And we think that that will foster more innovation and creativity.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  39. For example, we always had this prompt meaning of life. We test our models based on this prompt. So if you have meaning of life as your prompt, then do you want an image generated and a diffusion model to decide what the meaning of life is, or do you want a language model to kind of think and go back and forth and come up with a description of a scene that's explaining the meaning of life and that's kind of the context of JSON prompting is the intermediate representation that we think language models can describe images in that format and then imagination can happen. In general, we see a lot of editing happening in the field and that's the new frontier. So I don't think we should expect the interaction to be only through text or JSON, but it's a combination of JSON and image, if I were to make a guess.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  40. About that, And the Reddit was really lashing out as our engineers, and one of our people said, Oh, we might fix this. And they were like, we might. Oh. I'm sorry. But the fact is, the community needs to also read the documentation and bear with us. This model is only trained with JSON prompting. And you have to provide JSON with that particular structure for you to get good quality output. So I don't know if it's a fish or bug. We did have some safety built into the model, but that is also detecting incorrect prompts. So if we just give it a one-word prompt, then you get this image is blocked by safety image back, but that's because your prompt is not a well-specified JSON. Now we don't want people to write in JSON. We don't think that's a natural way of interacting with these models. But I do strongly believe that we need to use all the AI innovation to build the best image generation and editing models.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Saw a lot of JSON prompting in your technical blog, which is very unique. And as I was trying out model, it seems like it was translating the text, the prompt to a JSON representation with implicit structure. Do you think JSON is the representation for image models going forward, or do you think there's another representation there?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  42. We take images and we turn them to text using visual language models. The very first models we were training three, four years ago would be based on the alt text that you can find on internet. That is, each image on the internet may have an alt text field associated with it, which describes what's in the image. But the problem is the alt text is often very short or inaccurate. And what we do now is we train models to go from image to text. And in this case, image to text with detailed bounding box information, detailed element information, if we care about text and we really want to make sure all the text in the images correctly described. And then we go from text to image backward. So it's kind of interesting. We gather all the images from internet. Some of them may have all text, some of them may not have alt text, and then we use AI to go from image to text. And then we train another AI model to go.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  43. I would say a lot of it is really listing all the possible changes and very carefully tuning each element of the model and see what happens. Obviously we try to gather as much data as possible. One of the standard recipes in industry is that

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  44. Yeah, it's kind of difficult to exactly describe what resulted in such an amazing model. I think a lot of it is focus and evaluation. Evaluating image models is actually a very difficult thing to do. There are lots of benchmarks out there, but people look at them and they're like, okay, this doesn't correlate with pixel fidelity that I care about or realism. You don't really want novice users to judge the quality of these models because they may be looking at small monitors that aren't really adjusted for color accuracy. And all this cared so much about quality photorealism and again takes accuracy. So throughout training, we always measure takes accuracy and we applate very detailed changes to the model and data and see how that results in performance.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  45. Of the things Sa stood out to us, which is what the community has been chatting about is how there's new ways of processing data as you're training the model, which is like you kind of let the model learn what is a bounding box and how to do the layering and color palettes. They want to talk more about some of the innovations you had during the training process on what made this model so good with these differentiating features.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  46. Really stylized typography for logo, t shirt design, graphic design in general. And so we continue to push forward, I think our previous model wasn't really beating the state of the orient next generation, but we continue to focus on that and we had a bunch of research breakthroughs. And with this model, despite the fact that it's very tiny, the next generation is very, very accurate.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  47. I remember at the time, just a few people building these models, and the question was how can we differentiate? What's unique about our model? And we said, okay, next generation accurate text is something we have. And then we released it and we were really surprised. It's just so many people were so excited about text generation. And then we realized, oh, actually, that's the whole graphic design and storytelling industry. Text is a very important part of image generation. And that became a very important part of our brand. So if you search Idogram, people talk about the quality of typography, the quality of text we are known for.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  48. I don't know if you remember, but the very first model we released three years ago, and at the time image generation was synonymous with garbal text and there were memes about Dali 2 generating travel posters with incorrect city names, which is fun to look at.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  49. Amazing. And one of the other things I immediately noticed about the model was how you can render super long texts, like paragraphs of texts completely accurately, which you either give the model and the prompt or you ask the model to come up with something and it does it really well and it's super impressive, honestly, reaching the level of things like nano banana or GPT image with an open source model. Was that something you guys really focused on? And sort of why did you think that was important?

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source

  50. And if you look at our prompts, it's like thousands of words, each element in the image, where it is in the image, we have layout control bounding box and a number of elements. And that's one of the key innovations here that unlocks a lot of, again, design use cases because you clearly want font control, you want layout control, and this model is very versatile, allows you to really fix certain elements, fix positioning and control the image generation in every detail possible.

    2026-06-15 · a16z Podcast · AI, Design, and the Power of Open Models · IDENTIFIED FROM THE TRANSCRIPT · source