AI tools directory
DiffBIR: Image Restoration
A system that leverages previously trained text-to-image diffusion models for use in image restoration.*ColabPro https://colab.research.google.com/github/camenduru/DiffBIR-colab/blob/main/DiffBIR_colab.ipynb *Github https://github.com/XPixelGroup/DiffBIR *Paper https://arxiv.org/abs/2308.15070
VideoComposer
Imagen2Vídeo and Vídeo2Vídeo. *Web https://videocomposer.github.io/ *Demo https://modelscope.cn/studios/damo/I2VGen-XL-Demo/summary *ColabPro https://colab.research.google.com/github/camenduru/I2VGen-XL-colab/blob/main/I2VGen_XL_colab.ipynb *Github https://github.com/damo-vilab/videocomposer *Papper https://arxiv.org/abs/2306.02018
SeamlessM4T
A universal translator. It is designed to provide high quality translation, allowing people from different language communities to communicate effortlessly through voice and text. *Web https://github.com/facebookresearch/seamless_communication *Colab https://colab.research.google.com/github/camenduru/seamless-m4t-colab/blob/main/seamless_m4t_colab.ipynb *Demo https://seamless.metademolab.com/ *HuggingFace https://huggingface.co/spaces/facebook/seamless_m4t *Github https://github.com/facebookresearch/seamless_communication *Papper https://dl.fbaipublicfiles.com/seamless/seamless_m4t_paper.pdf
Magenta: A tool utilizing machine learning to aid in the creative process of art and music.
Magenta is an open-source project exploring the role of machine learning as a tool in the creative process. It offers a collection of music creativity tools built on open-source models, utilizing cutting-edge machine learning techniques for music generation.*Web https://magenta.tensorflow.org/ *Demo https://magenta.tensorflow.org/demos *Ableton https://magenta.tensorflow.org/studio *JS https://github.com/magenta/magenta-js *Github https://github.com/magenta *Manual https://magenta.tensorflow.org/studio#:~:text=Studio%20v1.0.-,TABLE%20OF%20CONTENTS,-Overview
Music To Image: A tool that turns your music into unique images
"Music To Image", developed by "fffiloni", has the ability to convert music into images, allowing users to visualize music in a new and creative way. It generates unique images based on the characteristics of the inputted music, offering an innovative visual experience that complements the music.*Web https://huggingface.co/spaces/fffiloni/Music-To-Image
TokenFlow: Edit your videos with text prompts
From an input video and a text indication, you can edit the style, objects or characters in your video.*Demo https://huggingface.co/spaces/weizmannscience/tokenflow *Web https://diffusion-tokenflow.github.io/ *Github https://github.com/omerbt/TokenFlow *Paper https://arxiv.org/abs/2307.10373
AnimateDiff: Create an animated image from a text prompt
This system proposes a technique for creating a time-coherent sequence of images to obtain an animation from the description of an image.*Demo https://huggingface.co/spaces/guoyww/AnimateDiff *Web https://animatediff.github.io/ *Github https://github.com/guoyww/animatediff/ *Paper https://arxiv.org/abs/2307.04725
Word-As-Image for Semantic Typography: Semantic Typography Generator
A system that creates a font with a graphic style associated with the concept of the words that are written with it.*Web: https://wordasimage.github.io/Word-As-Image-Page/*Demo: https://huggingface.co/spaces/SemanticTypography/Word-As-Image*Github: https://github.com/Shiriluz/Word-As-Image*Paper: https://arxiv.org/abs/2303.01818
AudioGen – Generate audio effects from text prompts
Audiogen is a text-to-sound conversion model, created from Audiocraft, a Meta pytorch library for deep learning research on audio generation. *Web https://audiocraft.metademolab.com/audiogen.html *Colab https://colab.research.google.com/github/camenduru/audiogen-colab/blob/main/audiogen_colab.ipynb *Github https://github.com/facebookresearch/audiocraft/blob/main/docs/AUDIOGEN.md *Paper https://arxiv.org/abs/2209.15352
Pix2Pix Video: Text-guided video style editing
An implementation of Pix2Pix applied to an image sequence, which you can use with your own videos. Because it processes each image independently, the resulting video shows shifts in style while preserving the morphology of the input image. *Demo https://huggingface.co/spaces/fffiloni/Pix2Pix-Video*Colab https://colab.research.google.com/github/camenduru/pix2pix-video-colab/blob/main/pix2pix-video-colab.ipynb#scrollTo=Cp1aDyeElG57 *Code https://huggingface.co/spaces/fffiloni/Pix2Pix-Video/blob/main/app.py
ControlNet: From a sketch to an image
A system that adds input conditions to diffusion image-generation models, allowing images to be generated from sketches, depth data or other images together with a descriptive phrase. *Colab https://colab.research.google.com/drive/1VRrDqT6xeETfMsfqYuCGhwdxcC2kLd2P?usp=sharing *Github https://github.com/lllyasviel/ControlNet *Paper https://arxiv.org/abs/2302.05543
BLIP-2: Text and image chat
A system that enables conversations based on the contents of an image. *Colab https://colab.research.google.com/github/salesforce/LAVIS/blob/main/projects/img2prompt-vqa/img2prompt_vqa.ipynb#scrollTo=7428ac2d *Github https://github.com/salesforce/LAVIS *Paper https://arxiv.org/abs/2301.12597
Hyperreel: 6DOF video player
An optimised way to play six-degree-of-freedom videos: videos in which you can move through the scene in three-dimensional space. *Web https://hyperreel.github.io/ *Github https://github.com/facebookresearch/hyperreel *Paper https://arxiv.org/abs/2301.02238
Live 3D: Manga character modelling and animation
*Colab https://colab.research.google.com/github/transpchan/Live3D-v2/blob/main/notebook.ipynb *Github https://github.com/transpchan/Live3D-v2/
Arcane, Disney and Archer: Stylised cartoon generator
The Stable diffusion image generator trained on three specific styles using the Dreambooth method. *Demo https://huggingface.co/nitrosocke/Nitro-Diffusion
Plug-and-Play Diffusion: Text-based image-to-image editing
Another way to perform this task is an implementation based on stable_diffusion that changes the appearance of an image from text prompts while preserving its original structure. It works by synthetically reconstructing the original image and describing its components, then modifying those components independently according to the text input. *Web https://pnp-diffusion.github.io/ *Demo https://huggingface.co/spaces/hysts/PnP-diffusion-features *Github https://github.com/MichalGeyer/plug-and-play *Paper https://arxiv.org/abs/2211.12572
instruct pix2pix: Editing images from text
An implementation based on stable_diffusion that changes the appearance of an image from text prompts while preserving its original structure. *Web https://www.timothybrooks.com/instruct-pix2pix/ *Demo https://huggingface.co/spaces/timbrooks/instruct-pix2pix *HuggingFace https://huggingface.co/timbrooks/instruct-pix2pix *Github https://github.com/timothybrooks/instruct-pix2pix *Paper https://arxiv.org/abs/2211.09800
3d diffusion: 3D model generation from a single image
This model synthesises a 3D model by predicting multiple views from the single perspective provided by an image. *Web https://3d-diffusion.github.io/ *Paper https://arxiv.org/abs/2210.04628
DDSP-VST: Neural synthesis tool that transforms and enriches your creative process with innovative sounds
DDSP-VST is a tool that allows you to experiment with new sounds and transform your creative process. You can use it like a typical virtual instrument, integrating it into your workflow with your favorite MIDI sources and effects. Additionally, it offers controls to dial in realistic sounds or explore a wide array of timbres that diverge from the original sound. You can also create your own models with its free web trainer, allowing for even more customization of your sound experience.*Web https://magenta.tensorflow.org/ddsp-vst
Stable diffusion: Text to image
A pretrained, open-source text-to-image generator. It was one of the systems that attracted the most attention in 2022 because it was freely available, released together with its training weights and ready to use. *Demo https://huggingface.co/spaces/stabilityai/stable-diffusion *Paid demo https://beta.dreamstudio.ai/dream *Colab https://colab.research.google.com/github/TheLastBen/fast-stable-diffusion/blob/main/fast_stable_diffusion_AUTOMATIC1111.ipynb *Mac https://diffusionbee.com/ *Pc https://nmkd.itch.io/t2i-gui *Web https://stability.ai/blog/stable-diffusion-public-release