本記事は**Geminiの出力をプロンプト工学で整理した業務ドラフト(未検証)**です。 # 多様な推論テンプレートでGRPOを安定化:Prompt Augmentationによる数学推論のスケーリング 【要点サマリ】 DeepSeek-R1で採用された**GRPO**の訓練初期における不安定さを、多様な推論テンプレートの注入によって解消。 – **課題**:単一の思考テンプレート(
graph TD
A["Original Prompt"] --> B{"Prompt Augmentation"}
B -->|"Template 1"| C1["Group Sample 1"]
B -->|"Template 2"| C2["Group Sample 2"]
B -->|"Template N"| Cn["Group Sample N"]
C1 & C2 & Cn --> D["Reward Calculation"]
D --> E["Group Relative Advantage"]
E --> F["Policy Update"]
import random
TEMPLATES = [
"Step-by-step reasoning:\n<thought>\n{query}\n</thought>",
"Analyze the problem first:\n<thinking>\n{query}\n</thinking>",
"Detailed logical derivation:\n<reasoning>\n{query}\n</reasoning>"
]
def get_augmented_batch(queries):
augmented_prompts = []
for query in queries:
# ランダムにテンプレートを選択し多様性を確保
template = random.choice(TEMPLATES)
augmented_prompts.append(template.format(query=query))
return augmented_prompts
# GRPOの訓練ループ内でこれを使用し、Groupごとの多様な応答を生成
# 生成された応答はそれぞれのテンプレートに従い、異なる推論構造を持つ
| 手法 | GSM8K (Acc) | MATH (Acc) | 報酬の収束速度 |
|---|---|---|---|
| Baseline GRPO | 78.2% | 42.5% | 低速(不安定) |
| GRPO + PA (提案) | 84.5% | 53.1% | 高速(安定) |
arXiv:2502.14857 [cs.LG] – “Prompt Augmentation Scales up GRPO: Stabilizing Reinforcement Learning for Mathematical Reasoning”
DeepSeek-R1 Technical Report (Reference for GRPO basics)

