一句話總結
為「跑一整天的大重構」而生,與 Muse Code 共訓、1M 窗口一口氣跑完,MCP 工具調度目前業界第一。
Meta Superintelligence Lab 第三代 Muse 系列,2026-08-05 發佈,coding-focused update。1M context 推理模型,與 Muse Glimmer 30B 同源但定位相反:Spark 1.2 是閉源旗艦(API only),Glimmer 是其開源蒸餾本地版。
基本資訊
| 項目 | 內容 |
|---|---|
| 開發者 | Meta Superintelligence Lab |
| 發佈日 | 2026-08-05(與 Muse Code beta 同日) |
| 家族 | Muse Spark 1.0 (2026-04-08) → 1.1 (2026-07) → 1.2 |
| License | Proprietary(閉源 API);開源權重「coming weeks」已宣佈、截至 2026-09-03 尚未釋出 |
| 參數量 | 未公開 |
| 架構 | 推理模型 mandatory thinking |
| Context | 1,048,576 tokens(1M) |
| 輸入模態 | text / image / video / audio / PDF |
| Reasoning effort | minimal / low / medium / high / xhigh,medium 預設 |
| API ID | muse-spark-1.2 / muse-spark-1.2-contributor |
| 可用管道 | Muse Code、Meta Model API、OpenRouter |
架構與訓練
- Co-training with Muse Code:從 day 1 把 Muse Code 放進訓練迴圈 — rejection-sampled harness trajectories、goal/compaction/subagent recipe 優化、直接整合 toolset。跨多 harness 訓練保留泛化能力。
- Self-improvement loop:用 Spark 1.1 生成更難的 coding 環境 + 指令模板 → 自評 → 篩高品質資料訓練 1.2。
- Long-horizon:whole-repo generation、大型端到端專案、auto-research;planning + goal conditioning + context compaction + async/parallel tool calls。
Muse Code(Beta)
| 特性 | 說明 |
|---|---|
| 安裝 | curl -fsSL https://dev.meta.ai/install.sh | bash → muse |
| Async Background Agents | 常駐背景 agent,自行決定何時回報主 agent |
| Event Log | 本地 JSONL,replay-exact,crash 後 muse resume 續跑 |
| Fan-out 隔離 | 每子 agent 獨立 git worktree,平行不衝突 |
| Bundled Skills | /plan /grill /taste /goal,皆 explicit-invocation |
Benchmarks
⚠️ 判讀陷阱:官方 TB 2.1 / DeepSWE 是「模型+自家 agent」系統分數,且 Spark 1.2 與 Muse Code 是 co-trained,換 harness 掉 2-3 點。
Meta 自測(5 attempts 平均 pass@1)
| Benchmark | 1.2 | 1.1 | Opus 5 | Terra |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 82.9% | 76.2% | 86.7% | 81.8% |
| DeepSWE 1.1 | 59.3% | 53.0% | 65.0% | 64.8% |
| Meta Internal (440 PR) | 70.6% | 68.3% | 79.4% | 65.4% |
| GDPVal-AA v2 Elo | 1631 | 1371 | 1852 | 1577 |
| MCP Atlas | 90.3% | 88.1% | 85.8% | — |
獨立第三方(Artificial Analysis)
| 指標 | 分數 | 備註 |
|---|---|---|
| Intelligence Index | 54 | #13/185,追平 Grok 4.5;1.1=51、1.0=43 |
| GDPVal Elo | 1631 | #5 overall,+260 vs 1.1 |
| TB 2.1 獨立 | 80% | vs 官方 82.9%(harness 效應 2.9 點) |
| Vals Index | 71.88% | #5/45,$0.69/test 最便宜 |
定價
| Model ID | Input | Cached | Output | 限制 |
|---|---|---|---|---|
| muse-spark-1.2-contributor | $0.10 | $0.002 | $0.20 | rolling 5h token 限流 |
| muse-spark-1.2 | $1.25 | $0.15 | $4.25 | pay-as-you-go |
Contributor 是目前最激進的 frontier-adjacent 定價;thinking tokens 計為 output。零資料保留可向 Meta sales 申請。
Hacker News 社群反應(2026-09-03)
來源:HN 當日 #1 Muse Spark 1.3 — 355 pts / 244 comments(含大量 1.2 回顧),對照 1.2 首發 333 pts / 266c。
正面
- 像工具而非意見領袖最獲共鳴:@superfrank 稱 1.2 knows its weaknesses, just acted like a tool — what I want 90%+ of the time,擔憂 1.3 為 benchmark 變 over-helpful。
- 性價比:@majerep best free model on OpenCode;@WASDx 逆向重寫老遊戲 quite good and fast;contributor 被評 20x discount(@jumploops)、practically free for hobbyists(@7734128)。
- 價格透明度:@jmward01 稱 first quantifiable number 能量化「拿你 tokens 訓練」的價值;@apodolny reasonable and transparent。
- Simon 鵜鶘 SVG:同 prompt 1.3 4.2¢/38s 畫質勝 1.2(better frame/wing/hat),xhigh 7.5¢/1m34s。
質疑
- 天花板仍落後 frontier:Fast and cheap but Terra felt much more capable(@Gecko4072);web app 出錯回 Claude/Kimi(@mromanuk);bug 上會 loop thought summaries(@WASDx)。
- benchmaxxing:But is the score really reflective?(@cbg0);terrible smart summaries(@bermudi)。
- 隱私:Extracted AWS keys from contributor models?(@0xbadcafebee);貢獻版易被拿去跑無訓練價值的 batch(@2001zhaozhao)。
- 鵜鶘測試已飽和:saturated and meaningless(@gpt5)。
- 生態遷移難:less than 6 months behind frontier, only SOTA would make devs switch(@wrsh07 / @gehsty)。
評價與定位
- 定位:非 frontier 王,是長程 coding agent 的性價比王(contributor + MCP 第一 + 1M 窗口)。
- 優勢:MCP 90.3% 第一、1M + compaction、event log / worktree 隔離最完整、$0.10/$0.20 顛覆預算。
- 弱勢:純 coding 增幅溫和(內部僅 +2.3)、SciCode 退步、harness 耦合、無開源權重、參數未公開。
- 風險:開源權重無日期/license,co-train 依賴是長期風險。
部署
curl -fsSL https://dev.meta.ai/install.sh | bash
muse
# /model to muse-spark-1.2-contributor
# /effort high
完整研究頁:models/meta-muse-spark-1.2.md(vault)