LLM Structured Output 入門:等 AI 齋出乾淨 JSON,唔再靠 regex 拆答案
JSON mode、function calling、json_schema 三層做法逐一講,附 Python 實戰代碼
重點整理
- 痛點:直接問 LLM 攞資料,佢會夠埋「好嘅,以下是…」呢啲廢話——寫程式想攞乾淨 JSON 就頭痛
- 三層做法:prompt 講明(最弱)→ JSON mode(強制合法 JSON)→ json_schema(連欄位都幫你鎖死)
- 實戰:用 Ollama 本地模型 + format schema,零 API 費由影評抽齣電影名/評分/一句總結
- pitfalls:溫度調低、schema 唔好太深、數字欄位要驗證——LLM 出 JSON 都會出錯
點解要 Structured Output?
你叫 LLM「將呢篇影評嘅評分同重點列出嚟」,佢會答:
好嘅!以下是這篇影評的重點:\n1. 評分:8.5 分…
人類睇冇問題,但程式要 parse 就想死——前綴、編號、全形標點全部要執。Structured output 就係令模型齋出合法 JSON,唔好講嘢。
三層做法(由弱到強)
| 層級 | 做法 | 可靠度 |
|---|---|---|
| 1. Prompt 講明 | 「請只用 JSON 回答」 | ❌ 時靈時唔靈 |
| 2. JSON mode | API 參數強制輸出合法 JSON | ✅ 一定合法,但欄位隨佢 |
| 3. json_schema | 畀埋 JSON Schema,欄位/類型都鎖死 | ✅✅ 最穩 |
實戰:Ollama 本地模型 + Schema
import requests
schema = {
"type": "object",
"properties": {
"movie": {"type": "string"},
"score": {"type": "number"},
"one_liner": {"type": "string"},
},
"required": ["movie", "score", "one_liner"],
}
r = requests.post("http://localhost:11434/api/chat", json={
"model": "qwen3:8b",
"messages": [{"role": "user", "content": "從呢段影評抽齣電影名、評分同一句總結:<評論全文>"}],
"format": schema, # Ollama 支援直接畀 JSON Schema
"stream": False,
"options": {"temperature": 0},
})
import json as J
print(J.loads(r.json()["message"]["content"]))
# {'movie': 'Dune: Part Two', 'score': 8.5, 'one_liner': '...'}
雲端 API 一樣有對應:OpenAI 用 response_format={"type": "json_schema", ...},Anthropic 就用 tool use 強制結構。
實用 tips
- 溫度調低(0~0.3):結構化任務唔需要創意
- Schema 唔好太深:巢狀超過三層,細模型好易漏欄位
- 數字要驗證:出嚟都係要 pydantic / 手動 check,唔好盲信
- enum 鎖選項:情感分類就用
"enum": ["pos", "neg", "neutral"],唔會出第四個答案
總結
Structured output 係「LLM 由玩具變工具」嗰一腳——出嘅嘢可以直接入資料庫、入 pipeline。記住三層做法:能唔能用 prompt 解決就 prompt,認真項目就一定上 json_schema。
