MediaGPT for Image Generation with Enhanced Text
Text-to-image models still fail at the thing media work depends on most: legible, correctly rendered text inside the image. MediaGPT is a generative AI system that treats language and image synthesis as one pipeline rather than two tools, pairing strong language models — ChatGPT and Phi-3 — with Stable Diffusion and ControlNet to drive intelligent text-image alignment on an interactive canvas. Layout can be customised freely while semantic coherence and aesthetic balance hold.
The paper positions the system for designers, educators, marketers and content creators: effortless multi-modal integration behind easy-to-use tools, so producing high-quality, contextually appropriate content stops being a two-pass job. It was presented at ICT4SD 2025 and published by Springer Nature in Lecture Notes in Networks and Systems.
