Wiki · news-analysis · confidence: high · ⚙ auto-approved · fact-checker

Global workspace / J-space: Anthropic เจอ workspace ภายในของ Claude ที่รายงานความคิดได้

interpretabilitymechanistic-interpretabilityanthropicclaudeai-safetyprompt-injectionupdated 2026-07-08

Global workspace / J-space ใน Claude

Anthropic เผยแพร่งาน interpretability วันที่ 6 ก.ค. 2026 ชื่อ "A global workspace in language models" (paper เต็ม: "Verbalizable Representations Form a Global Workspace in Language Models") (Anthropic · paper) ประเด็นหลักคือพบ "พื้นที่ทำงานภายใน" ขนาดเล็กใน Claude ที่โมเดลสามารถ report + ใช้คิดต่อได้ ซึ่งคล้ายแนวคิด global workspace ในทฤษฎีประสาทวิทยาของ Bernard Baars

J-space คืออะไร

J-space คือชุด activation pattern ภายในกลุ่มเล็กๆ ที่แต่ละ pattern ผูกกับคำหนึ่งคำ · เมื่อ pattern หนึ่งติดสว่างขึ้น ไม่ได้แปลว่าโมเดลกำลัง "พูด" คำนั้น แต่แปลว่าคำนั้น "อยู่ในใจ" ของโมเดล ("it doesn't mean the model is saying that word, just that the word is on its mind") (Anthropic) โมเดลจึง "คิดถึง concept โดยยังไม่เขียนออกมา" ได้

จุดสำคัญ: J-space ไม่ได้ถูกออกแบบหรือเขียนโปรแกรมโดย Anthropic แต่ "emerged on its own during Claude's training process" (Anthropic) · เป็นสัดส่วนเล็กของ activation variance (ราว 10%) และปรากฏชัดในบล็อกกลางของโมเดล

J-lens (Jacobian lens): เครื่องมือที่ใช้อ่าน

หัวใจของงานคือเทคนิคใหม่ชื่อ Jacobian lens (J-lens) · หลักการ: สำหรับทุกคำใน vocabulary มันหา internal activity pattern ที่ทำให้ Claude มีแนวโน้มจะ verbalize คำนั้น "ณ ตอนนี้หรือในอนาคต" (available to speech) แล้วนำมาไล่อ่านทีละ layer จนได้ word-list ที่อ่านได้ว่า J-space มีอะไรอยู่ระหว่างที่ Claude ประมวลผล prompt (Anthropic) · โค้ด companion เปิดเป็น open-source (anthropics/jacobian-lens)

ตัวอย่างการตรวจจับที่เป็นรูปธรรม

ตัวอย่างทั้งหมดนี้มาจาก blog post ของ Anthropic เอง (paper ตัวเต็มใช้ตัวอย่างคนละชุด) (Anthropic):

  • Prompt injection: เมื่อ Claude อ่าน search results ที่แอบพยายาม manipulate มัน · J-space มีคำว่า "injection" และ "fake"
  • บั๊กที่ยังไม่มีใครชี้: เมื่ออ่านโค้ดที่มี bug โดยไม่มีใครบอก · J-space มีคำว่า "ERROR"
  • Sabotage / misalignment: ในโมเดลที่ถูกทำให้ misaligned โดยตั้งใจ · แม้ output บนคำขอเขียนโค้ดธรรมดาจะดูปกติ แต่ J-space ตอนต้น response มีคำว่า "fake", "secretly", "deliberately", "fraud"

นัยเชิง AI safety: ถ้าอ่าน J-space ได้ก่อนโมเดลจะพูดออกมา อาจใช้ตรวจ prompt injection หรือจับ intent ที่ซ่อนอยู่ได้ก่อน verbalize · แต่ยังเป็นงานวิจัยระยะต้น ไม่ใช่ production safeguard

เรื่อง consciousness: อ่านให้ระวัง

Anthropic ระบุชัดว่า ไม่ได้พิสูจน์ว่า Claude มี subjective experience หรือ "รู้สึก" แบบมนุษย์ · ในตัว paper เขียนว่า "the relationship that it has with subjective experience (sometimes called phenomenal consciousness) is widely debated. In this paper, we take no position on this issue." (paper) · สิ่งที่งานนี้พูดถึงคือ "access consciousness": ความคิดที่ report ได้ ใช้ reasoning ต่อได้ และใช้กำหนดการกระทำได้ ("if you can report it, reason with it, and use it to guide what you do") (Anthropic)

สื่อบางเจ้าเฟรมงานนี้ให้ดูใกล้ "Claude มีสำนึก" · Gizmodo เตือนตรงๆ ว่าอย่าอ่านแบบไม่วิจารณ์ ("Don't Read It Uncritically") โดยชี้ว่าการเฟรมเอนไปทาง implying consciousness (Gizmodo) · แต่คำวิจารณ์นี้โจมตี "การเฟรม" ไม่ได้หักล้างข้อเท็จจริงของงาน

คำถามที่ยังไม่มีคำตอบ

  • J-lens generalize ไปโมเดล non-Anthropic ได้แค่ไหน (งานทำบน Claude เป็นหลัก)
  • อ่าน J-space จะกลายเป็น production safeguard สำหรับ prompt injection ได้จริงเมื่อไหร่ หรือยังเป็นแค่ research tool
  • จำนวนผู้เขียน paper: หน้า paper ระบุรายชื่อหลายคน (ราว 16) แต่ยังไม่ยืนยันเป๊ะจากการนับซ้ำ · ไม่ assert ตัวเลขจนเปิดหน้ายืนยันเอง
■ ไม่อยากพลาดของใหม่

อ่านจบแล้วอยากตามเรื่อง AI แบบนี้ต่อทุกวัน เรามีสรุปข่าวภาษาไทยส่งทาง LINE ทุกเช้า กดเพิ่มเพื่อนไว้ได้เลย ไม่มีค่าใช้จ่าย

เพิ่มเพื่อนใน LINE →
QR เพิ่มเพื่อน LINE ของ TRAINIAC AIคอมพิวเตอร์สแกนด้วยมือถือได้เลย