小模型都玩得轉 Function Calling:本地 LLM Agent 工具呼叫實戰指南
7B 級 model 點樣穩定呼叫工具——schema 設計、retry 策略、Ollama structured output 實測心得
重點整理
- 痛點:大 model function calling 好穩,但本地 7B 小模型一遇到複雜 tool 就亂嚟——點解?點救?
- 三大關鍵:tool 數量控制、schema 極簡化、錯誤 feedback loop——每個都附本地 Ollama 實測對比
- 進階技巧:structured output(JSON schema)模式強制格式、多工具 fallback 設計、 hallucinated argument 偵測
- 附完整可抄 code:Python + Ollama 極簡 agent loop,80 行行齊 tool calling 全流程
小模型 function calling 真係得嗎?
好多人覺得 function calling 係大 model 專利。但 2026 年嘅本地小模型(Qwen、Llama 系 7B 級)經過 fine-tune 之後,工具呼叫其實已經好能打——前提係你要識得遷就佢。
大 model 可以一次過食 20 個工具都揀得啱;小模型俾佢 20 個工具,佢會揀錯、串錯參數名、甚至發明一個唔存在嘅工具。解法唔係換 model,係改設計。
三大關鍵設計
1. Tool 數量:一次過最多 5 個
實測(Qwen3 8B,Ollama):5 個工具以內,揀中率 >90%;去到 10 個就跌穿 70%。
解法:工具分組 + 兩段式呼叫。第一輪先問 model「你想用邊一類工具」(search / file / calculator),第二輪先俾該類別嘅具體工具。
2. Schema 極簡化
小模型對嵌套 JSON schema 好敏感。呢個 schema 大 model 食得落:
{
"name": "search_web",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "search keywords"},
"filters": {"type": "object", "properties": {"date_range": {"type": "string"}}}
}
}
}
但小模型會將 filters 搞到亂晒。改做:
{
"name": "search_web",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string"},
"days_back": {"type": "integer"}
}
}
}
原則:flat(無嵌套)、參數名用一個英文字、每個參數一定要有 description。小模型好靠 description 猜你想點。
3. 錯誤 Feedback Loop
小模型第一次輸出錯 JSON 好正常。唔好即刻報錯收工,將錯誤訊息塞返入對話俾佢重試:
result = call_model(messages, tools)
for attempt in range(3):
try:
args = json.loads(result.arguments)
break
except json.JSONDecodeError as e:
messages.append({"role": "tool", "content": f"Invalid JSON: {e}. Please output valid JSON matching the schema exactly."})
result = call_model(messages, tools)
實測:一次 retry 之後成功率由 ~75% 升到 ~93%。兩次 retry 去到 97%。
Structured Output 模式:強制唔使估
Ollama 支持 format 參數直接俾 JSON schema,model 一定會輸出合 schema 嘅 JSON(constrained decoding,唔係靠提示詞求佢):
import json, requests
schema = {
"type": "object",
"properties": {
"tool": {"type": "string", "enum": ["search_web", "read_file", "none"]},
"query": {"type": "string"}
},
"required": ["tool", "query"]
}
resp = requests.post("http://localhost:11434/api/chat", json={
"model": "qwen3:8b",
"messages": [{"role": "user", "content": "搵下今日香港天氣"}],
"format": schema,
"stream": False,
})
print(resp.json()["message"]["content"])
呢個方法嘅好處:100% 合 schema,壞處:model 冇得「答唔到所以唔呼叫」——所以 enum 入面要加一個 none 值俾佢婉拒。
極簡 Agent Loop(可抄)
import json, requests
TOOLS = {
"get_weather": lambda city: f"{city} 今日 28°C,多雲",
"calc": lambda expr: str(eval(expr)), # demo only
}
def run(user_input):
messages = [{"role": "user", "content": user_input}]
for _ in range(5): # 最多 5 步,防止無限循環
resp = requests.post("http://localhost:11434/api/chat", json={
"model": "qwen3:8b",
"messages": messages,
"tools": [
{"type": "function", "function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {"type": "object", "properties": {
"city": {"type": "string", "description": "city name"}}}},
},
{"type": "function", "function": {
"name": "calc",
"description": "Evaluate a math expression",
"parameters": {"type": "object", "properties": {
"expr": {"type": "string", "description": "e.g. 2+2*10"}}}},
},
],
"stream": False,
}).json()
msg = resp["message"]
messages.append(msg)
if not msg.get("tool_calls"):
return msg["content"] # 冇工具呼叫 = 最終答案
for tc in msg["tool_calls"]:
fn = tc["function"]["name"]
args = tc["function"]["arguments"]
out = TOOLS.get(fn, lambda **k: "unknown tool")(**args)
messages.append({"role": "tool", "content": json.dumps(out, ensure_ascii=False)})
return "超過步數上限,已中止。"
print(run("倫敦而家幾多度?如果加 10 度又係幾多?"))
常見坑
- Hallucinated arguments:model 會用 schema 冇嘅參數名。解法:收到 arguments 之後按 schema 白名單過濾,多餘嘅刪走唔好報錯。
- 中文工具描述反而差:實測工具 name/description 用英文、user prompt 用中文,揀中率最高。
- temperature 調低:tool calling 場景設 0–0.3,創意係敵人。
- 唔好俾 model 自由選「要先做咩」:流程控制寫死喺你個 loop 度,model 只負責揀工具同填參數。
總結
小模型 function calling 嘅公式:≤5 個工具 + flat schema + description 齊全 + JSON 錯誤 retry + structured output 兜底。全部做齊,7B 級本地 model 都可以穩定跑一個實用 agent——零 API 費、數據不出機。
延伸閱讀
- Structured Output / JSON mode 入門——今篇手法嘅理論基礎
- Ollama vs llama.cpp 實戰對決——未裝好環境嘅先睇呢篇
