- calendar_today August 10, 2025
OpenAI launched “Images in ChatGPT,” which brings a groundbreaking feature enabling users to generate images directly inside the ChatGPT interface. The newly released GPT-4o model powers this breakthrough, which lets users generate images during their chats while representing a major advancement in AI content creation.
“Images in ChatGPT” extends its advanced image generation capabilities to all ChatGPT users on Plus, Pro, Team, and free subscription plans. OpenAI spokesperson Taya Christianson explained that free tier users will follow similar usage restrictions as DALL-E 3, which allows three image generations daily, but these limits might change depending on demand. DALL-E fans receive confirmation of ongoing access through a specialized GPT.
OpenAI research lead Gabriel Goh described GPT-4o as a groundbreaking “omnimodal” platform that processes multiple data types like text and visual/audio content. The model now boasts improved “binding” features, which resolve frequent AI image generation issues. The GPT-4o model excels at preserving object attribute relationships for 15 to 20 objects without confusing their colors or shapes, unlike earlier models that encounter this issue.
The model’s enhanced text rendering capabilities stand out as a key improvement. The generation of text within AI images has historically resulted in confusing or incoherent characters. According to Goh, the development process required numerous iterations over many months to perfect. The team has achieved text rendering consistency that makes text in images functionally usable despite the ongoing challenges with perfect text rendering for small elements.
The system’s architecture does not use diffusion models like most image generators, but instead uses an autoregressive method. The sequential image generation process, moving left to right followed by top to bottom, offers improved text rendering and binding capabilities because it mirrors the text generation approach.
The briefing demonstrated OpenAI’s system capabilities through creating scientific diagrams with precise labeling of Newton’s prism experiment, alongside multi-panel comics with consistent dialogues and informational posters with accurate text. The demonstration included practical examples that showed how the system can create transparent background images for items such as stickers and restaurant menus, as well as logos.
As multimodal product lead at ChatGPT, Jackie Shannon highlighted the system’s capability to utilize world knowledge. In my image creation process, I face my own skill constraints, yet I apply the comprehensive world knowledge I’ve accumulated. Through its access to world knowledge, the model allows users to obtain images of Newton’s prism experiment without providing detailed descriptions.
OpenAI acknowledges that image generation takes slightly more time now, but maintains that the improved quality and capabilities make the additional wait worthwhile. Shannon acknowledged the need for latency improvements yet emphasized that superior image quality and comprehensive world knowledge compensate for the extra waiting time.
Safeguards and User Ownership: Ensuring Responsible AI Image Generation
OpenAI took steps to address misuse concerns by implementing strong protective measures. The system functions to block CSAM requests and inhibit both watermark removal and creation of sexual deepfakes. All images produced by the system will contain standard C2PA metadata, which identifies them as creations of OpenAI despite lacking visual watermarks. The company runs its own verification tools for images.
Shannon acknowledged that while no system can be flawless for this purpose, they keep advancing their protective measures and consider the current state as their foundation. Users who generate images with ChatGPT retain ownership of these images and can use them according to the platform’s usage rules in any manner they choose.
OpenAI’s “Images in ChatGPT” feature enhances its main product functionality and expands AI creativity boundaries by giving users a potent visual expression instrument through their chat interface. OpenAI shows its dedication to enhancing user experience through this new feature while actively addressing risks related to advanced AI image generation technology. The efforts to enhance binding capabilities alongside text rendering while integrating protective measures demonstrate a commitment to producing a tool that delivers strong performance alongside ethical responsibility.





