AI Research

Peer-reviewed work from the ATX Labs team.

MediaGPT for Image Generation with Enhanced Text

Text-to-image models still fail at the thing media work depends on most: legible, correctly rendered text inside the image. MediaGPT is a generative AI system that treats language and image synthesis as one pipeline rather than two tools, pairing strong language models — ChatGPT and Phi-3 — with Stable Diffusion and ControlNet to drive intelligent text-image alignment on an interactive canvas. Layout can be customised freely while semantic coherence and aesthetic balance hold.

The paper positions the system for designers, educators, marketers and content creators: effortless multi-modal integration behind easy-to-use tools, so producing high-quality, contextually appropriate content stops being a two-pass job. It was presented at ICT4SD 2025 and published by Springer Nature in Lecture Notes in Networks and Systems.

Open full sizeView on SpringerLinkFirst page · full text via Springer
First page of “MediaGPT for Image Generation with Enhanced Text”, showing the title, authors, abstract and keywords.

Working on something adjacent?

We collaborate with research teams and publish the results — tell us what you're exploring.