Recent progress in artificial intelligence (AI) does not seem to be slowing down, showcasing new technological capabilities at a rate that makes it difficult to imagine what may next significantly impact daily lives and businesses or what may become obsolete. Tech companies such as OpenAI, Google, Microsoft, Meta, NVIDIA, Apple, and Adobe are investing heavily in research and development (R&D) and adopting emerging AI capabilities into their products as quickly as possible to avoid missing any major disruptive opportunities.
Current-day discussion on AI mostly revolves around generative AI technologies and transformer-based products such as ChatGPT or GPT-4. These technologies are being incorporated into search engines like Bing (bing.com), while huge amounts of AI-generated images from diffusion models are circulating on social media. Generative AI for media synthesis can refer to the generation and manipulation of images, videos, audio, and even 3D content. Media can be generated with (text, image, audio, etc.) or without input (random noise).
The commercialization activities and opportunities of AI are limitless and span almost every industry sector (telecommunication, health, transportation, education, energy, entertainment, etc.). This paper focuses on:
the commercialization of major generative AI technologies for media synthesis in recent history (what is real and what is not); and the new technological capabilities and potential disruptive products on the horizon.
Deepfakes: Still Harmless Synthetic Media
Since the emergence of so-called deepfake technologies, public concerns have been raised about potential dangers in the context of disinformation, fraud, and harassment, especially due to the technology’s wide adoption for non-consensual internet pornography. Deepfake technologies are a form of media synthesis, often based on generative AI that is focused on manipulating specific human subjects in a video (e.g., through face swapping, facial puppeteering, or lip-synching). The risk of deepfake technologies being used for malicious purposes seems particularly tangible as they can generate extremely convincing results (traditionally only achievable by professional visual effects studios), can make people appear to say and do anything one wishes them to, and the technology is largely accessible to anyone as it is relatively easy to learn and requires no technical expertise.
While technologies for media manipulation have continued to progress at a steady rate, deepfakes themselves seem to be rather harmless and not as catastrophic as many experts predicted. Deepfakes have been used for online fraud and harassment cases, during political elections, or by Russian activists in the Ukraine War, but so far have been ineffective as a tool for disinformation (especially compared to fake news in general). Many of the deepfakes (images or videos) circulating on the internet are generated by non-experts, and despite looking impressive, they still appear artificial and do not require sophisticated detection algorithms to be uncovered.
Productization of Generative Media
Mobile Apps, Augmented Reality Filters, and Social Media Videos
Beyond some gimmicky and entertaining mobile apps and augmented reality (AR) filters (Snap, TikTok, etc.), deepfake technologies may not initially seem to have any deeper meaningful purpose beyond manipulating videos and adding new visual effects to social media posts. Some of the most common effects include inserting a user’s face into a video clip (e.g., Zao, ReFace) by uploading a single photo and selecting a pre-curated video, and swapping the face of a user video with that of a celebrity (e.g., Impressions.ai) by uploading a video and selecting a pre-trained face model of a celebrity. Facial reenactment using a technique called first order motion has also gained popularity, where an arbitrary portrait of a person is uploaded and immediately reenacted, with users being able to create viral videos of politicians singing or reenacting people from the past (e.g., DeepNostalgia, MyHeritage).
The demand for new and more impressive tools for self-expression are driving researchers in both academia and industry to continue to push the limits of media synthesis capabilities (e.g., higher resolution, real-time performances, more control, less artifacts, more accessibility, etc). This has resulted in new, sophisticated filters such as de-aging, gender swaps, cartoon filters, and the introduction of new algorithms with innovative input modalities (single image reenactment, text-based specifications, photo-real avatars from videos, etc.).
Virtual Assistants, Marketing Videos, and Universal Translators
While essential in many entertainment applications, other commercial sectors have also explored the use of digital humans to improve, automate, and scale their services through the use of generative AI. Several companies have developed human-like virtual assistant solutions (e.g., Soul Machines, Uneeq) but they fail to appeal to customers due to their “uncanny valley” appearances (the feeling of discomfort caused by viewing imperfect computer-generated faces)Footnote 118Footnote 119 . Despite technological advances in the application of graphics engines (e.g., Epic Games / MetaHumans, unrealengine.com) or generative AI to enhance photorealism in those avatars (e.g., Samsung Neon, Pinscreen, etc.), virtual assistants still struggle to replace real humans. They currently lack sophisticated responses, and their voices and facial expressions are often absent of emotion and empathy.
However, due to recent advancements in large language models (LLM) such as ChatGPT and emerging research in motion synthesis, the mass adoption of highly convincing and realistic human-like AI agents may be closer at hand (within two to three years), especially if they have the ability to interact in real-time. In the meantime, several startup companies (e.g., Synthesia, Colossyan, etc.) are exploring the use of generated videos of pre-recorded humans in a non-interactive setting where marketing and training videos are generated at scale for enterprise applications. An actor and/or voice can be selected through a web interface, and a text script provided as input in order to generate video content automatically on a server. These solutions typically use a text-to-speech solution (e.g., third party or proprietary where voices can be customized), and a speech-to-face video generator that uses an audio input and video frames as training data (e.g., for Synthesia: ten minutes of an actor performing a speech in a well-lit studio environment and facing the camera head-on).
These methods are more advanced than the popular wav2lip algorithm and generate higher resolution and better quality results. Similar technologies have also been adopted by Chinese Tencent and Korean companies such as DeepBrain in the context of generating news anchors and marketing material at scale. Tencent for instance charges only $145 USD for each subject (either half or full body) and supports both English and Chinese languages. Despite their high level of fidelity, the resulting human performance generated from speech still appears slightly robotic during conversations, and mass adoption is still limited.
Google recently announced at their I/O Conference an enterprise-level service called Universal Translator, which allows educational content creators to translate their videos into multiple languages. The solution uses translated voices as input to generate synchronized lip movements in the videos. The translated voice input is also produced using a generative translation model that mimics the voice and tone of the speaker but in a different language. As of now, this offering is only available for select and authorized content creators (e.g., partnership with Arizona State University), which can assist in preventing its use for malicious applications.
Cheaper and Faster: Visual Effects (VFX) and Visual Dubbing for Hollywood
Whether it is to create digital stunt doubles, bring deceased stars back to life, or de-age an older actor, computer-generated digital actors are widely deployed in some of the most memorable blockbuster films (e.g., Star Wars, Furious 7, Terminator: Dark Fate, and The Curious Case of Benjamin Button). However, these effects typically rely on sophisticated visual effects studios (e.g., Industrial Light & Magic, Weta Digital, MPC, Framestore, etc.), cost millions of dollars, and take months of work to produce a few seconds of footage. Visual effects related to human facial performances are particularly expensive and difficult to achieve due to the “uncanny valley” effect.
As open source deepfake solutions (such as faceswap-GAN and Deep Face Lab) emerged and became freely accessible on the Internet, hobbyists and deepfake artists began generating entertaining videos by swapping celebrities in short video clips. While it was possible to produce highly convincing deepfakes, the resolution was often still too poor for film production. However, these methods quickly caught the attention of VFX producers as a tool for enhancing their conventional VFX pipelines in order to save cost and impact story telling. Visual effects companies such as Industrial Light & Magic (ILM) have explored the use of deepfake technologies for de-aging actors (e.g., Mark Hamill in Star Wars, Harrison Ford in Indiana Jones 5). This is achieved by replacing the faces of aged actors or doubles with neural renders built from younger footage of the same actor and combining them with 3D models and video compositing techniques.
AI startup companies such as Pinscreen and Metaphysic provide complete AI visual effects solutions for face replacement in film production. Metaphysic is known for its viral Tom Cruise deepfakes circulating on TikTok and their recent Elvis face replacement on America’s Got Talent.
Pinscreen innovated the development of a number of GAN-based neural face rendering technologies (most notably PaGAN, “photoreal avatar GAN”), which were originally developed to enhance the realism of 3D avatars for interactive 3D and metaverse applications. In 2022, the company started to shift its focus in the VFX space through a partnership with Netflix and Amazon Studios, and launched on a number of high profile TV shows (e.g., The Manifest), blockbuster movies (e.g., Slumberland 2022), and advertisements (Nike, Balenciaga, etc.) using generative AI technologies. AI VFX services include end-to-end processing for face replacement, facial reenactment, aging/de-aging, and visual dubbing. Pinscreen’s key advantage consists of being able to handle very short cinematic shots and deliver high-fidelity 4K HDR output, allowing for the processing of close-up shots, extreme side views, and dramatic/dynamic lighting conditions. The process requires specialized GAN-based data augmentation and AI enhancement procedures to generate unseen data from sparse views collected from film footage and improved architectures for high resolution and temporally coherent video synthesis.
Despite the growing demand of AI VFX services such as face replacement, aging, and de-aging, these remain relatively niche applications and are highly show-dependent. One scalable market is in visual dubbing for films and TV shows, where foreign films can be watched in any desired language while also having actors’ lip movements perfectly synchronized to speech. Feature films are much harder to process than video clips that are captured in controlled settings (such as of news anchors, marketing and training materials) due to the complexity of scenes, lack of training data, and the extremely high quality requirements in cinema (4K HDR).
In 2022, Pinscreen became the world’s first company to fully lip sync a full feature film using its proprietary generative AI pipeline, demonstrated on the film The Champion – translated from German/Polish to English. The process combined state-of-the-art generative AI and an integrated VFX pipeline, which allowed Pinscreen to complete the processing of a 90 minute film in less than three months. The approach can handle an existing movie, and only requires additional video recordings of the voice actors during the dubbing process. Other players in AI VFX such as Flawless.ai are trying to enter the market for visual dubbing but have limited technical capabilities as they only offer speech-to-face reenactment as opposed to video performance as input. They have demonstrated some visual dubbing examples on select video clips, but not entire movies.
Diffusion-Based Text-To-Image Generation
With breakthroughs such as OpenAI’s Dall-E and recent advancements in diffusion and transformer-based models such as Stable Diffusion, image generation capabilities that outperform traditional GAN-based methods in terms of image quality, resolution, and diversity are now possible. The latter property is particularly significant as it enables highly effective text-to-image generation, where users can input an arbitrary text prompt allowing the model to generate an image that reflects this prompt accurately. Incorporating text input is typically enabled by using a CLIP-encoder that can map the prompt into a text embedding which is then used as a condition for generating an image using a progressive de-noising process (the generator), which is typically based on iteratively using a deep neural network based on a U-net architecture for image-to-image translation.
While training those models is easier and more reliable than GANs, diffusion model training is extremely resource intensive, typically requiring weeks of training and hundreds of high performance GPUs (A100s). Consequentially, those models are often trained by companies who have large GPU resources (e.g., OpenAI, Stability.ai, Google, etc.), while labs in academia and smaller companies rely on pre-trained models they can further fine-tune. The latest and most popular commercial solutions include OpenAI’s Dall-E-2, Midjourney (via Bot on Discord), as well as Stability.ai’s solution (available as web interface Dream Studio) and APIs. While incredibly realistic images can be generated, they are still prone to noticeable artifacts, and production-level fine control is not yet possible. Some level of control through scribbles or abstract skeletons have been recently demonstrated (e.g., ControlNet), but the generated images always come with unpredictable details and appearances. As a result, diffusion-based methods are not yet suitable for production-quality video generation as they lack controllability and temporal consistency.
Summary and Future Capabilities
Generative AI capabilities for media synthesis (images, video, audio) are constantly evolving. Generated image qualities are improving (e.g., higher resolution, less artifacts, and more semantically realistic results) and are more diverse, enabling natural text prompts as input. Similar to when GANs were introduced, the research community is focusing on enabling better controllability, more predictable outputs, and temporally consistent generations for videos, as well as the ability to handle other modalities such as neural 3D content. Due to this technology’s accessibility and performance in generating convincing content, concerns around its potential misuse have been raised by the public. So far, these media synthesis technologies and deepfakes have not been extensively weaponized, even though they pose a potential threat.
In the coming years, society is expected to witness further technological breakthroughs in generative AI. These breakthroughs will enable new commercial opportunities, including general online video generation services (e.g., a YouTube that can take any text prompt as input and generate the desired video on-the-fly), real-time and fully interactive videos (e.g., advertisements that can interact with a viewer in real-time), as well as fully immersive and photorealistic AI-generated environments for Metaverse applications. With the recent announcements of new augmented reality/virtual reality (AR/VR) headsets such as Apple’s Vision Pro and Meta’s MetaQuest 3, it is foreseeable that the demand for sophisticated 3D content will grow and generative AI will play a key role in enabling content creation.