AI News · 8 ข่าว

AI News · 2026-07-08

Models

Meta ปล่อย Muse Image: image model สาย sub-agent architecture · #2 บน Arena รองเฉพาะ GPT Image 2

Meta เปิด Muse Image (8 ก.ค.) เป็นโมเดล generate ภาพที่ใช้สถาปัตยกรรม sub-agent ทำงานร่วมกับโมเดลหลัก Muse Spark (แยก agent สำหรับ run code + ค้นข้อมูล) · ทำได้ทั้ง text-to-image, สร้าง QR code/GIF, แก้ภาพเฉพาะจุดโดยคงรายละเอียดเดิม, ค้นข้อมูลปัจจุบันมาช่วยออกแบบ · Meta เคลมว่าติดอันดับ 2 บน Arena ทั้งหมวด text-to-image, single-image edit, multi-image edit "รองเฉพาะ GPT Image 2" · เปิดใช้บน Meta AI, Instagram Stories (US), WhatsApp (บางประเทศ) · Muse Video ยังเป็น preview และยังมีปัญหา sync เสียง-ภาพ + ฟิสิกส์การเคลื่อนไหวไม่สมจริง

ทำไมต้องรู้ · image-gen ยักษ์ใหญ่ขยับอีกตัว กระทบคอร์ส presentation/marketing โดยตรง · จุดน่าสอนคือ sub-agent architecture (โมเดล gen ภาพเรียก agent ย่อยไปทำงานเฉพาะทาง) เป็น pattern เดียวกับ agentic workflow ที่สอนอยู่ · และเป็น data point ล่าสุดว่าคำเคลม "ตัวท็อปของโลก" ของ image/video tool ตัวเดียวล้าสมัยเร็ว (Muse เองยังเป็นรอง GPT Image 2)

Tools

Claude Code 2.1.203: ชุดแก้ reliability ของ background agent (ต่อเนื่องจาก 2026-07-07)

ขยับจาก 2.1.202 · เป็น maintenance release ล้วน ไม่มี breaking UI ใหม่: เพิ่มเตือน login ใกล้หมดอายุ, MCP roots/list ส่ง working directory ของ session ได้, แก้ background agent ค้างบน macOS (regression จาก 2.1.196), แก้ background session ไม่ตอบสนองเมื่อ daemon token หมดอายุ, แก้ subagent work ติดค้างเวลากลับเข้า claude agents, ปรับ context-usage performance, แก้ PATH inheritance ของ background agent บน Windows

ทำไมต้องรู้ · เวอร์ชันที่ใช้สอน workshop ขยับต่อ แต่ไม่มี demo ไหนพังเพิ่มจาก 2.1.200 (permission Manual ยังเป็น ■ เดิม F04) · ประเด็น background agent reliability + Windows PATH ตรงกับผู้เรียนสาย dev ที่รัน agent ยาวๆ

Simon Willison ปล่อย sqlite-utils 4.0 stable: database migrations + nested transactions (ต่อเนื่องจาก 2026-07-07)

หลังไล่ RC มาหลายตัว (rc2 07-05 · rc3/rc4 07-07) release 4.0 ตัวจริงออกแล้ว (124th release) · ของใหม่: ระบบ schema migration, nested transaction ผ่าน db.atomic(), รองรับ compound foreign key · Fable 5 ช่วยไล่บั๊กสำคัญก่อน stable · พร้อมกันปล่อย sqlite-migrate 0.2 ที่ยุบ library แยกตัวเก่าเข้ามาเป็น compatibility shim อิงระบบ migration ใหม่

ทำไมต้องรู้ · ปิด thread AI-assisted release + cross-model review ของสัปดาห์ (RC → stable) · เคสจริงที่เล่าเป็น QA workflow ได้: ให้โมเดลรีวิว release ก่อนออก stable ไม่ใช่ปล่อยดิบ · log ปิดเรื่องไว้กันเล่าซ้ำ

Research

Anthropic เจอ "global workspace / J-space" ใน Claude: อ่านคำที่โมเดลคิดถึงแต่ยังไม่พูดได้ ⚙

Anthropic เผยแพร่งาน interpretability (6 ก.ค. · paper "Verbalizable Representations Form a Global Workspace in Language Models") · พบ J-space = ชุด activation เล็กๆ ที่แต่ละ pattern ผูกกับคำหนึ่งคำ เมื่อสว่างแปลว่าคำนั้น "อยู่ในใจ" โมเดล (ยังไม่พูด) · เครื่องมือใหม่ J-lens (Jacobian lens) อ่าน word-list นี้ได้ทีละ layer · ตัวอย่างตรวจจับ (จาก blog): อ่าน search ที่แอบ manipulate → J-space มี "injection"/"fake" · โค้ดมีบั๊กที่ยังไม่ชี้ → "ERROR" · โมเดลที่ถูกทำให้ misaligned บนงานปกติ → "fake/secretly/deliberately/fraud" · Anthropic ย้ำว่าไม่ได้พิสูจน์ว่า Claude มี subjective experience และ take no position เรื่อง phenomenal consciousness · พูดถึงแค่ "access consciousness"

ทำไมต้องรู้ · หลักฐานสดว่า mechanistic interpretability อ่าน "สิ่งที่โมเดลกำลังคิดก่อนพูด" ได้บางส่วน · เป็นทั้งมุมบวก AI trust/safety (ต่อ thread trust ของเดือน · เชื่อม prompt-injection detection) และเคสสอน media literacy ชั้นดี: หัวข้อที่ "ดูเหมือน consciousness" ถูกสื่อขยายเกินจริงทันที ต้องแยก access vs phenomenal · ผ่าน fact-checker gate (ทุก claim ยืนยันจากหน้า Anthropic เอง · confidence high) → เก็บเข้า wiki

AutomationBench-AA: benchmark อิสระ (AA + Zapier) วัด agent ทำงาน SaaS จริง · ทุกโมเดลยังแหก business rule ⚙

Artificial Analysis + Zapier เปิด leaderboard อิสระ (Zapier สร้าง benchmark · AA รัน/โฮสต์เอง) วัด agent ทำงานข้ามแอป SaaS จริง · รอบ AA-run ประเมิน 657 tasks บน 40 simulated environment (Gmail, Sheets, Slack, Salesforce, Zendesk, Jira, HubSpot) 6 domain · headline score = สัดส่วน objective ที่ทำได้โดยไม่ละเมิด guardrail (business rule) ไม่ใช่ full-task success · ผล: Fable 5 48.6% · Opus 4.8 48.5 · Gemini 3.5 Flash 42.6 · GPT-5.5 xhigh 42.1 · GLM-5.2 max 27.8 (open ดีสุด) · ทุกโมเดลละเมิด guardrail (0.46 ครั้ง/task Gemini ต่ำสุด ถึง 1.26 Qwen3.7 Plus) · Finance ยากสุด (ทำได้ราวครึ่งของ Support/Operations) · ระวัง: ถ้าใช้ metric เข้มแบบ full-task ครบถูกต้อง paper ระบุ frontier ดีสุดยังต่ำกว่า 10%

ทำไมต้องรู้ · เลข third-party ที่วัด agent ในงานธุรกิจจริง (ต่างจาก vendor เคลมเอง คู่กับ Hy3) · ตอบคำถามลูกค้าคอร์ส Cowork/agent ตรงๆ ว่า "ปล่อย agent ทำงานข้ามระบบแทนคนได้จริงไหม" ด้วยตัวเลข: ท็อปทำ objective โดยไม่แหกกฎได้ไม่ถึงครึ่ง + ทุกตัวยังละเมิด business rule = human-in-the-loop + guardrail ยังจำเป็น · ผ่าน fact-checker gate (numeric claim cross-verified · attribute scope + นิยาม metric ชัด) → เก็บเข้า wiki

Agent memory thread: A-TMA แก้ fact เก่าชนใหม่ + ReContext เล่น evidence ก่อน generate (ต่อเนื่องจาก 2026-07-04)

สองเปเปอร์จากฉบับ smol.ai 07-06 · A-TMA (Ghost Memory): จัดการ conflict ระหว่าง fact เก่ากับ fact ปัจจุบันใน assistant ที่รันยาว · เคลม +0.240 absolute บน LTP benchmark เมื่อต่อเข้ากับ Graphiti · ReContext: harness แบบ training-free ที่ replay evidence จากภายในโมเดลก่อน generate เพื่อใช้ context ยาวได้ดีขึ้น (ทดสอบบน 8 dataset ขนาด 128K)

ทำไมต้องรู้ · ต่อ thread agent memory ของสัปดาห์ (คู่กับ AutoMem 07-04, wiki-structured memory 07-03) · ตรงหลักการ Learning OS ของเรา (distill ความรู้ ไม่เก็บ transcript ดิบ) · แต่ทั้งคู่ยัง cite ผ่าน tweet ยังไม่มี paper URL ตรง เลย log เป็น thread ต่อเนื่อง ไม่ promote

Long-context papers ใหม่บน arXiv: HiLS attention (extrapolate 64x) + KV cache compression (KVpop/SeKV)

arXiv RSS 07-07 · "Hierarchical Sparse Attention Done Right" (2607.02980) เสนอ HiLS attention ที่ extrapolate context ได้เกิน 64 เท่าของความยาวตอน train · คู่กับกลุ่ม KV cache compression: KVpop (2607.05061) ใช้ learned eviction policy จาก future-attention supervision ลด memory · SeKV (2606.31145) จัด semantic span ข้าม GPU-CPU ลด memory ~53%

ทำไมต้องรู้ · พื้นฐานเชิงกลไกว่าทำไมโมเดลรุ่นใหม่เสิร์ฟ context ยาวได้ถูกลง (ต่อ thread FlashMorph 07-05, Drowning in Documents 07-04) · ใช้ตอบคำถาม "1M context ทำได้ยังไงไม่ให้แพงระเบิด" ในคลาสสาย engineer · ยัง early-stage log ไว้เป็น thread

Business

YC CEO Garry Tan เคลมส่งโค้ด AI 37,000 บรรทัด/วัน · dev ไปแกะเว็บพบ bloat หนัก (HN 106 pts)

Garry Tan (CEO Y Combinator) เคลมว่าส่งโค้ดที่ AI ช่วยเขียน ~37,000 บรรทัด/วัน · วิศวกรโปแลนด์ (ประสบการณ์ 13 ปี) ไปแกะเว็บของเขา (GStack) พบปัญหาหนัก: 169 request รวม 6.42 MB (เทียบ Hacker News 7 request 12 KB), ส่งไฟล์ test 28 ไฟล์ให้ผู้ใช้ทุกคน, preload JS controller ที่ไม่ได้ใช้ 78 ตัว, PNG ไม่บีบอัด (บางไฟล์ 2 MB), CSS ว่าง, เนื้อหาซ้ำ · โหลดข้อมูลมากกว่าเว็บ minimal ~500 เท่า · consensus บน HN: "ทุกบรรทัดคือ liability ไม่ใช่ productivity"

ทำไมต้องรู้ · ต่อ thread "AI confidence theater" (Elena Verna 07-04) + "วัด adoption ด้วย metric ผิด" (Meta tokenmaxxing 07-03) · เคสสอนผู้บริหารตรงๆ ว่า lines-of-code = liability ไม่ใช่ output · อวด "AI เขียนโค้ดวันละหมื่นบรรทัด" ต้องถามต่อว่าโค้ดนั้นคุณภาพเท่าไร (คู่กับ Building to the Test 07-03: agent ส่งงานตามที่ตรวจ ไม่ใช่ตามที่ขอ)

■ ไม่อยากพลาดของใหม่

สรุปข่าวแบบนี้ระบบทำทุกเช้า ถ้าอยากให้คัดเรื่องเด่นส่งถึงมือทาง LINE โดยไม่ต้องแวะมาเช็คเอง กดเพิ่มเพื่อนไว้ได้เลย ไม่มีค่าใช้จ่าย

เพิ่มเพื่อนใน LINE →
QR เพิ่มเพื่อน LINE ของ TRAINIAC AIคอมพิวเตอร์สแกนด้วยมือถือได้เลย