Meta Muse Spark 1.2 研究:1M 視窗、MCP 90.3% 與最透明的 Contributor 定價(含 HN 社群反應)

Muse Spark 1.2 是 Meta 的 coding-focused update,1M context、與 Muse Code 共訓,MCP Atlas 90.3% 業界第一,contributor /bin/bash.10//bin/bash.20 定價最激進。HN 2026-09-03 登頂 355pts,社群肯定 tool-like 特質但質疑 benchmaxxing 與 harness 耦合。

一句話總結

為「跑一整天的大重構」而生,與 Muse Code 共訓、1M 窗口一口氣跑完,MCP 工具調度目前業界第一。

Meta Superintelligence Lab 第三代 Muse 系列,2026-08-05 發佈,coding-focused update。1M context 推理模型,與 Muse Glimmer 30B 同源但定位相反:Spark 1.2 是閉源旗艦(API only),Glimmer 是其開源蒸餾本地版。

基本資訊

項目 內容
開發者 Meta Superintelligence Lab
發佈日 2026-08-05(與 Muse Code beta 同日)
家族 Muse Spark 1.0 (2026-04-08) → 1.1 (2026-07) → 1.2
License Proprietary(閉源 API);開源權重「coming weeks」已宣佈、截至 2026-09-03 尚未釋出
參數量 未公開
架構 推理模型 mandatory thinking
Context 1,048,576 tokens(1M)
輸入模態 text / image / video / audio / PDF
Reasoning effort minimal / low / medium / high / xhigh,medium 預設
API ID muse-spark-1.2 / muse-spark-1.2-contributor
可用管道 Muse Code、Meta Model API、OpenRouter

架構與訓練

  • Co-training with Muse Code:從 day 1 把 Muse Code 放進訓練迴圈 — rejection-sampled harness trajectories、goal/compaction/subagent recipe 優化、直接整合 toolset。跨多 harness 訓練保留泛化能力。
  • Self-improvement loop:用 Spark 1.1 生成更難的 coding 環境 + 指令模板 → 自評 → 篩高品質資料訓練 1.2。
  • Long-horizon:whole-repo generation、大型端到端專案、auto-research;planning + goal conditioning + context compaction + async/parallel tool calls。

Muse Code(Beta)

特性 說明
安裝 curl -fsSL https://dev.meta.ai/install.sh | bashmuse
Async Background Agents 常駐背景 agent,自行決定何時回報主 agent
Event Log 本地 JSONL,replay-exact,crash 後 muse resume 續跑
Fan-out 隔離 每子 agent 獨立 git worktree,平行不衝突
Bundled Skills /plan /grill /taste /goal,皆 explicit-invocation

Benchmarks

⚠️ 判讀陷阱:官方 TB 2.1 / DeepSWE 是「模型+自家 agent」系統分數,且 Spark 1.2 與 Muse Code 是 co-trained,換 harness 掉 2-3 點。

Meta 自測(5 attempts 平均 pass@1)

Benchmark 1.2 1.1 Opus 5 Terra
Terminal-Bench 2.1 82.9% 76.2% 86.7% 81.8%
DeepSWE 1.1 59.3% 53.0% 65.0% 64.8%
Meta Internal (440 PR) 70.6% 68.3% 79.4% 65.4%
GDPVal-AA v2 Elo 1631 1371 1852 1577
MCP Atlas 90.3% 88.1% 85.8%

獨立第三方(Artificial Analysis)

指標 分數 備註
Intelligence Index 54 #13/185,追平 Grok 4.5;1.1=51、1.0=43
GDPVal Elo 1631 #5 overall,+260 vs 1.1
TB 2.1 獨立 80% vs 官方 82.9%(harness 效應 2.9 點)
Vals Index 71.88% #5/45,$0.69/test 最便宜

定價

Model ID Input Cached Output 限制
muse-spark-1.2-contributor $0.10 $0.002 $0.20 rolling 5h token 限流
muse-spark-1.2 $1.25 $0.15 $4.25 pay-as-you-go

Contributor 是目前最激進的 frontier-adjacent 定價;thinking tokens 計為 output。零資料保留可向 Meta sales 申請。

Hacker News 社群反應(2026-09-03)

來源:HN 當日 #1 Muse Spark 1.3 — 355 pts / 244 comments(含大量 1.2 回顧),對照 1.2 首發 333 pts / 266c。

正面

  • 像工具而非意見領袖最獲共鳴:@superfrank 稱 1.2 knows its weaknesses, just acted like a tool — what I want 90%+ of the time,擔憂 1.3 為 benchmark 變 over-helpful。
  • 性價比:@majerep best free model on OpenCode;@WASDx 逆向重寫老遊戲 quite good and fast;contributor 被評 20x discount(@jumploops)、practically free for hobbyists(@7734128)。
  • 價格透明度:@jmward01 稱 first quantifiable number 能量化「拿你 tokens 訓練」的價值;@apodolny reasonable and transparent
  • Simon 鵜鶘 SVG:同 prompt 1.3 4.2¢/38s 畫質勝 1.2(better frame/wing/hat),xhigh 7.5¢/1m34s。

質疑

  • 天花板仍落後 frontier:Fast and cheap but Terra felt much more capable(@Gecko4072);web app 出錯回 Claude/Kimi(@mromanuk);bug 上會 loop thought summaries(@WASDx)。
  • benchmaxxing:But is the score really reflective?(@cbg0);terrible smart summaries(@bermudi)。
  • 隱私:Extracted AWS keys from contributor models?(@0xbadcafebee);貢獻版易被拿去跑無訓練價值的 batch(@2001zhaozhao)。
  • 鵜鶘測試已飽和:saturated and meaningless(@gpt5)。
  • 生態遷移難:less than 6 months behind frontier, only SOTA would make devs switch(@wrsh07 / @gehsty)。

評價與定位

  • 定位:非 frontier 王,是長程 coding agent 的性價比王(contributor + MCP 第一 + 1M 窗口)。
  • 優勢:MCP 90.3% 第一、1M + compaction、event log / worktree 隔離最完整、$0.10/$0.20 顛覆預算。
  • 弱勢:純 coding 增幅溫和(內部僅 +2.3)、SciCode 退步、harness 耦合、無開源權重、參數未公開。
  • 風險:開源權重無日期/license,co-train 依賴是長期風險。

部署

curl -fsSL https://dev.meta.ai/install.sh | bash
muse
# /model to muse-spark-1.2-contributor
# /effort high

完整研究頁:models/meta-muse-spark-1.2.md(vault)

發佈留言

發佈留言必須填寫的電子郵件地址不會公開。 必填欄位標示為 *