Meta launched Muse Image on July 7, 2026 — its first image-generation model from Meta Superintelligence Labs (MSL), and the division's second major release after Muse Spark in April. Unlike conventional generators that map a prompt straight to pixels, Muse Image works as an agent: it searches the web, writes and runs code, refines its own output, and scales quality with extra test-time compute. Alongside it, Meta previewed Muse Video, its first video model, built on the same pretraining base and notable for generating native, synchronized audio — something most text-to-video models still don't do. Here's what actually shipped, what's only a preview, how the models rank, and the privacy controversy that arrived with them.
Table of Contents
TL;DR — what people are asking
| Question | Answer |
|---|---|
| Can I use it now? | Muse Image, yes — free in the Meta AI app, meta.ai, Instagram Stories (US), and WhatsApp (limited countries). Muse Video is preview-only for now. |
| Is it #1 on Arena? | No. Muse Image ranks No. 2 in text-to-image, single-image edit, and multi-image edit (Elo, July 5, 2026). Muse Video is No. 3 in text-to-video. |
| What makes it "agentic"? | It uses web search for facts and trends, executes code for plots, QR codes, and HTML games, self-refines its own drafts, and improves with more test-time compute. |
| How does it relate to Muse Spark? | They share tools. Spark and Muse Image can co-plan for GIFs, websites with embedded images, and interactive media. |
| Provenance? | Content Seal — an invisible watermark on every generation, plus a preview detector tool. Video support coming. |
| vs OpenAI / Google? | Meta claims strong editing and multi-reference composition. Arena scores measure human preference, not task accuracy — test on your own workflows. |
What is Muse Image?
Muse Image — reportedly codenamed Mango during development — is Meta's first fully in-house image model from Superintelligence Labs, the AI division led by Alexandr Wang. Meta describes it as a model that "uses advanced reasoning to understand complex prompts, seamlessly blending multiple photos into high-quality creations you can download and share anywhere."
Meta's own framing captures what's different about the architecture:
Instead of directly mapping prompts to images, Muse Image operates as an agent: it invokes search and coding tools to improve accuracy, self-refines its own generations, and improves through scaling test-time compute.
In practice, that means the model can plan before it draws. The headline features at launch: text rendering inside images, photo editing through markup (circle a thing, describe the change), AI-generated infographics, room redesigns using real products from the web or Facebook Marketplace, preset templates for quick creation, and — most controversially — the ability to incorporate public Instagram content into personalized visuals.
Tool use: coding, search & commerce
The agentic loop is the story here. Meta demonstrated three tool families:
| Tool | What Meta demonstrates |
|---|---|
| Coding | RL-taught code execution for accurate plots, scannable QR codes, and conditioning on rendered figures. Paired with Muse Spark: animated GIFs, websites with embedded images, and interactive visual games. |
| Search | Web search for current events, product catalogs, and scientific diagrams. Meta's internal ablation shows a higher win rate with search enabled. |
| Commerce | A Facebook Marketplace workflow that restyles a room using real, purchasable listing references (US). |
Think of it as loop engineering applied to pixels: plan → tool → draft → verify → revise. If the model needs a formula, a real product, or a chart with correct numbers, it doesn't hallucinate them — it fetches or computes them first.
Self-refinement: emergent, not scripted
Notably, Meta did not hand-design a fixed "critique then redraw" template. Self-refinement emerged during reinforcement learning because revised images scored higher reward. The model learned to make local edits for small errors, fully regenerate when the composition fails, or pivot to tools for factual tasks — Meta's demo thread shows a magazine spread where the model corrected its own formula notation mid-generation.
Test-time compute: quality as a slider
Meta reports approximately log-linear Elo gains as combined text-reasoning and visual-generation compute increases. Three product-relevant findings from the post:
- Best-of-N saturates quickly. Blindly sampling more images stops helping fast.
- Deliberate reasoning plus tool calls scales better than blind multi-sampling.
- Reasoning and tools compound — search fills knowledge gaps that reasoning alone can't.
For builders routing media APIs, the takeaway is to treat inference budget as a user-facing quality slider, not a hidden cost center — the same lesson agent harnesses learned for text.
Muse Video: the preview with native audio
Muse Video, built on the same pretraining base as Muse Image, was previewed on the same day but isn't publicly available yet. Its standout claim is native audio support — the model generates synchronized sound alongside the visuals, where most competing text-to-video models still output silent clips you have to score in post. Meta says Muse Video will power an upcoming AI video generation platform and eventually extend to its apps the same way Muse Image has.
If you need working video generation today rather than a preview, our AI video tools comparison and Sora vs Runway head-to-head cover what's actually shippable right now.
How it ranks on Arena
Strong debut, not a takeover — on Arena's human-preference Elo leaderboards (as of July 5, 2026):
- Muse Image: No. 2 in text-to-image, single-image edit, and multi-image edit.
- Muse Video: No. 3 in text-to-video — behind Google's gemini-omni-flash (1527 Elo) and ByteDance's Seedance 2.0 (1482), ahead of the rest of the field.
One caveat worth repeating: Arena measures which output people prefer at a glance, not task accuracy. A model can win preference votes while failing your specific edit workflow. If precise editing or multi-reference composition is your use case, run your own comparisons.
Where you can try it
| Surface | Status |
|---|---|
| Meta AI app & meta.ai | Available now, free |
| Instagram Stories | Available (US) |
| WhatsApp chats | Available (limited countries) |
| Facebook & Messenger | Rolling out gradually |
| Advantage+ (advertisers) | In the coming weeks |
| Muse Video | Preview only |
Content Seal & provenance
Every Muse Image generation carries Content Seal, an invisible watermark, and Meta is previewing a detection tool that checks whether an image carries one. Video support is planned. It's a reasonable step, but the usual limits apply: invisible watermarks help platforms and researchers flag content at scale, and they can degrade under heavy re-editing, screenshots, or recompression. Treat it as a signal, not proof.
The privacy controversy
Muse Image launched into immediate criticism over one feature: inside the Meta AI app, users can @ mention a public Instagram account to pull that account's photos into an AI generation. In other words, Muse Image can put an Instagram user into someone else's AI creation without that user's awareness.
Meta's position is that the feature only works with public content, but critics note that "public photos" and "consent to be AI-generated" are very different things — especially for images involving realistic composites of real people. If you have a public Instagram account and don't want your photos used this way, watch for the account-level controls Meta says are coming, or set your account to private. We'll update this section as Meta responds to the backlash.
Why Meta built this: the business angle
Muse Image isn't a research flex — it's an engagement and ads play. On Meta's Q1 2026 earnings call, management framed the Muse family (starting with Muse Spark) as the foundation for personal and business AI agents across its apps, and reported double-digit increases in Meta AI sessions per user after Spark's rollout.
The monetization math explains the urgency. Advertising made up nearly 98% of Meta's Q1 2026 revenue ($55.02 billion, up 33% year over year). More than 8 million advertisers already use at least one of Meta's generative AI creative tools, and advertisers using its AI video features saw over 3% higher conversion rates in testing. Folding Muse Image into Advantage+ creative lowers the ad-production barrier for small businesses — more usable creative, more inventory, more spend.
The competitive pressure is real, too: Alphabet's Search & Other revenue grew 19% year over year in Q1 2026 with revenue from generative-AI-model products up nearly 800%, while Amazon's ad business grew 24% to $17.2 billion in the quarter. Keeping image and video creation inside Meta's own apps — instead of losing those sessions to third-party generators — is the strategic point of the entire Muse line.
Frequently asked questions
Yes — Muse Image is free in the Meta AI app, on meta.ai, in Instagram Stories (US), and in WhatsApp in limited countries. Advertiser access via Advantage+ creative is rolling out separately.
It's agentic: rather than mapping a prompt directly to an image, it can search the web for facts, write and execute code for accurate plots and QR codes, refine its own drafts, and use more test-time compute for higher quality.
If your account is public, yes — users can @ mention public Instagram accounts in the Meta AI app to include their photos in generations, without notifying the account. Setting your account to private prevents this.
Muse Video is preview-only for now. Meta says it will power an upcoming AI video platform; no public release date has been announced. Its headline feature is native, synchronized audio generation.
Every generation carries Content Seal, an invisible watermark, and Meta is previewing a detector tool. Watermarks can weaken under heavy editing or recompression, so treat detection as a strong signal rather than absolute proof.
This article is independent editorial content based on Meta's announcement, Arena leaderboard data, and public reporting as of July 9, 2026. Velkar AI has no paid or affiliate relationship with Meta, OpenAI, Google, or ByteDance. Details of preview products can change — check Meta's official announcements for current availability. See our Affiliate Disclosure for our general policy.