Multimodal AI

News, repos, and tools about Multimodal AI. Auto-updated every 6 hours.

Latest News
ChatGPT Sketch turns your bad drawings into detailed AI images
The Verge · Sep 8, 2026
Roland launches Melody Flip AI plugin for generative music creation with 250 genre palettes
The Verge · Sep 4, 2026
Stability AI, creator of Stable Diffusion image generator, secures $76 million in new funding
TechCrunch · Aug 25, 2026
MindTopo benchmark reveals gaps in VLM spatial and topological reasoning
Microsoft Research · Aug 12, 2026
Major YouTube creators are facing backlash for accepting AI money
The Verge · Aug 21, 2026
Robin Williams’ Instagram account brought back to fight ‘AI abuse’
The Verge · Aug 18, 2026
Firefox Smart Window adds real-time web search and natural language browsing history features
The Verge · Aug 18, 2026
Why people aren’t buying Mark Zuckerberg’s AI future
TechCrunch · Aug 16, 2026
Woman joins lawsuit alleging stepfather used Grok to create explicit images from childhood photo
TechCrunch · Aug 15, 2026
Google now lets you disable Gemini's visible watermarks
The Verge · Aug 14, 2026
Google to let users remove visible watermarks from AI-generated images, videos, songs
TechCrunch · Aug 14, 2026
Meta releases Muse Glimmer: open-source 30B multimodal model for local agentic use
Hugging Face · Aug 10, 2026
Suno plans watermarks for AI-generated music to meet industry standards
Ars Technica · Aug 6, 2026
How we built a realtime system for responsive voice AI in six months
OpenAI Blog · Aug 3, 2026
Google Earth AI image generator sparked deepfake concerns before rollback
The Verge · Jul 31, 2026
Google Earth’s AI deepfake tool only lasted one day
The Verge · Jul 31, 2026
Canadian legislator reads out apparent LLM prompt instruction in floor speech
Ars Technica · Jul 24, 2026
Meta's Content Seal AI watermarking criticized for duplicating Google's SynthID
The Verge · Jul 22, 2026
Meta removes AI image feature from Instagram following backlash
TechCrunch · Jul 10, 2026
Google's deepfake detector system debunks McConnell hoax image
TechCrunch · Jul 8, 2026
Trending Repos
Asabeneh/30-Days-Of-Python
The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days.
Python · 73.3k stars
ultralytics/ultralytics
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classificatio
Python · 61.5k stars
freestylefly/awesome-gpt-image-2
Prompt as Code | GPT Image 2 / 2.5 提示词与案例库,530+ 个案例、20+ 套工业级模板与可复用 Skills,新增 2.5 同提示词对比专区,附完整提示词与生成记录,持续更新。
JavaScript · 30.6k stars
heygen-com/hyperframes
Write HTML. Render video. Built for agents.
TypeScript · 48.6k stars
huggingface/speech-to-speech
Build voice agents with open-source models
Python · 13.1k stars
microsoft/Ontology-Playground
Free, open-source web app for learning about ontologies and Microsoft Fabric IQ. Explore a catalogue of pre-built ontolo
TypeScript · 2.6k stars
huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and
Python · 165.1k stars
calesthio/OpenMontage
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and pr
Python · 57.0k stars
mudler/LocalAI
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU requir
Go · 49.0k stars
oobabooga/textgen
Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.
Python · 47.6k stars
ToolJet/ToolJet
Open-source foundation of ToolJet AI - the enterprise app generation platform for internal tools, dashboards, business a
JavaScript · 40.9k stars
bytedance/UI-TARS-desktop
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
TypeScript · 38.9k stars
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
Python · 35.7k stars
labring/FastGPT
FastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities suc
TypeScript · 29.6k stars
deepset-ai/haystack
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modula
Python · 26.5k stars
budtmo/docker-android
Android in docker solution with noVNC supported, video recording and mcp server
Python · 15.8k stars
HBAI-Ltd/Toonflow-app
Toonflow 是开源一站式 AI 短剧创作工具,将小说、剧本快速转化为动画短剧。集成 AI 编剧、智能分镜、角色与视频生成,跨平台桌面端轻量部署,助力创作者低成本批量产出视觉内容。Toonflow is an open-source A
TypeScript · 15.4k stars
FlowiseAI/Flowise
Build AI Agents, Visually
TypeScript · 55.5k stars
modelcontextprotocol/inspector
Visual testing tool for MCP servers
TypeScript · 10.8k stars
gorse-io/gorse
AI powered open source recommender system engine supports classical/LLM rankers and multimodal content via embedding
Go · 9.8k stars
Related Tools
Midjourney
Leading AI art generator known for highly aesthetic, photorealistic, and artistic image outputs. Web app and Discord.
DALL-E
OpenAI image generator integrated into ChatGPT, creating and editing images from natural language prompts.
Stable Diffusion
Open-source AI image model by Stability AI that can run locally or via API for unrestricted generation.
Leonardo
AI image and video generation platform with fine-tuned models for game assets, design, and art.
Ideogram
AI image generator known for exceptional text rendering accuracy within generated images.
Flux
State-of-the-art open-source image generation model by Black Forest Labs with excellent prompt adherence.
Runway
Leading AI video generation platform with Gen-4 Turbo model for creating and editing cinematic video from text and images.
Pika
AI video creation platform that generates and edits video clips from text prompts with cinematic quality.
Kling
Advanced AI video generation model producing high-quality, physics-aware video from text and image inputs.
Sora
OpenAI text-to-video model capable of generating realistic scenes with complex motion and camera movements.
Other Topics