docs(10_Wiki): 위키 전체 재구성 — Topic_* 폴더를 4개 카테고리로 통합 + 대규모 중복 제거
Topic_Agent/Topic_Blog/Topics/Topics_Biz/Topics_Meeting/Topics_Rag의 마크다운 지식 문서를 Topic_General/Topic_Programming/Topic_Graphic/Topic_Business 4개 카테고리로 재분류. - 중복 제거: frontmatter의 status:duplicate/merged + duplicate_of/redirect_to 필드로 자기 자신을 중복으로 선언한 리다이렉트 stub 1032개 제거, 완전 동일 내용 파일 472개 제거, 동일 파일명·다른 내용 충돌 시 더 큰(완전한) 버전만 유지(162개 제거) — 총 1639개 중복 제거. - 분류: 폴더 단위로 명확한 항목(AI_and_ML/Coding/Architecture 등 → Programming, Comfyui/Visual_Effects → Graphic, Topics_Biz/Topics_Meeting/사업 등 → Business, Poetic_Blog_Writing/창의성/Game_Design 등 → General)은 폴더 우선순위로, 나머지 혼재 폴더(Topic_Agent/Topic_Blog/Topics 루트/Thinking & Reasoning/Other/UI_UX_Assets)는 title/tags 키워드 스코어링으로 파일 단위 분류(불명확한 경우 General로 폴백). 원본 폴더명은 "From_*" 서브폴더로 보존해 추적 가능성 유지. - 최종 배치: Programming 2784 / General 1608 / Graphic 285 / Business 249 = 4926개 문서. - 에이전트 운영 상태(.astra/.agent/.obsidian/sessions/memory/_company/docs/lessons/_shared/src)는 지식 콘텐츠가 아니므로 재분류 대상에서 제외하고 원위치 유지. - Topics/Topic_email(상위 보호 폴더 Topic_email과 파일명 100% 중복) 삭제 — 보호 폴더 자체는 미변경. - 완전히 비게 된 Topic_Agent/Topic_Blog/Topics_Biz/Topics_Rag 폴더 제거.
This commit is contained in:
@@ -0,0 +1,166 @@
|
||||
---
|
||||
id: ai-llm-eval-patterns
|
||||
title: LLM Evaluation — Golden Set / LLM-as-Judge / 회귀
|
||||
category: Coding
|
||||
status: draft
|
||||
source_trust_level: B
|
||||
verification_status: conceptual
|
||||
created_at: 2026-05-09
|
||||
updated_at: 2026-05-09
|
||||
tags: [ai, llm, eval, testing, vibe-coding]
|
||||
tech_stack: { language: "TS / Python", applicable_to: ["Backend"] }
|
||||
applied_in: []
|
||||
aliases: [LLM eval, golden dataset, LLM-as-judge, regression, Promptfoo, Braintrust]
|
||||
---
|
||||
|
||||
# LLM Evaluation
|
||||
|
||||
> "느낌상 좋아짐" 은 측정 X. **golden dataset + 자동 채점**. Prompt 변경 / 모델 변경 시 회귀 검출. Promptfoo / Braintrust / LangSmith.
|
||||
|
||||
## 📖 핵심 개념
|
||||
- Golden set: input + expected output 쌍.
|
||||
- Metric: exact match / similarity / structured / LLM-as-judge.
|
||||
- Eval = unit test for LLM. 매 PR 마다 실행.
|
||||
- LLM-as-judge: 정답이 자유 형식일 때 다른 LLM 이 채점.
|
||||
|
||||
## 💻 코드 패턴
|
||||
|
||||
### 단순 자체 eval
|
||||
```ts
|
||||
const cases = [
|
||||
{ input: '2+2', expected: '4' },
|
||||
{ input: 'capital of France', expected: 'Paris' },
|
||||
];
|
||||
|
||||
let pass = 0;
|
||||
for (const c of cases) {
|
||||
const out = await callLLM(c.input);
|
||||
if (out.includes(c.expected)) pass++;
|
||||
else console.log('FAIL', c.input, '→', out);
|
||||
}
|
||||
console.log(`${pass}/${cases.length}`);
|
||||
```
|
||||
|
||||
### Promptfoo (yaml)
|
||||
```yaml
|
||||
# promptfooconfig.yaml
|
||||
prompts:
|
||||
- "Answer concisely: {{question}}"
|
||||
|
||||
providers:
|
||||
- openai:gpt-4o-mini
|
||||
- openai:gpt-4o
|
||||
- anthropic:claude-haiku-4-5
|
||||
|
||||
tests:
|
||||
- vars: { question: "Capital of France?" }
|
||||
assert:
|
||||
- type: contains
|
||||
value: "Paris"
|
||||
- type: latency
|
||||
threshold: 2000
|
||||
- type: cost
|
||||
threshold: 0.001
|
||||
|
||||
- vars: { question: "Bank vault security tips" }
|
||||
assert:
|
||||
- type: llm-rubric
|
||||
value: "Lists at least 3 security measures, mentions surveillance"
|
||||
```
|
||||
|
||||
```bash
|
||||
promptfoo eval
|
||||
```
|
||||
|
||||
### LLM-as-judge
|
||||
```ts
|
||||
async function judge(input: string, output: string, criteria: string): Promise<{score: number, reason: string}> {
|
||||
const r = await openai.chat.completions.create({
|
||||
model: 'gpt-4o',
|
||||
messages: [
|
||||
{ role: 'system', content: 'You are a strict evaluator. Score 0-5. Output JSON: {"score":N,"reason":"..."}' },
|
||||
{ role: 'user', content: `Input: ${input}\nOutput: ${output}\nCriteria: ${criteria}` },
|
||||
],
|
||||
response_format: { type: 'json_object' },
|
||||
});
|
||||
return JSON.parse(r.choices[0].message.content!);
|
||||
}
|
||||
```
|
||||
|
||||
### Pairwise comparison (A vs B)
|
||||
```ts
|
||||
// 실험: 두 prompt 결과 — 어느 게 나은지
|
||||
async function pairwise(input: string, outA: string, outB: string) {
|
||||
const r = await openai.chat.completions.create({
|
||||
model: 'gpt-4o',
|
||||
messages: [{ role: 'user', content: `Compare A and B for "${input}".\nA: ${outA}\nB: ${outB}\nWhich is better and why? JSON: {"winner":"A"|"B"|"tie","reason":"..."}` }],
|
||||
response_format: { type: 'json_object' },
|
||||
});
|
||||
return JSON.parse(r.choices[0].message.content!);
|
||||
}
|
||||
```
|
||||
|
||||
### Structured output 검증
|
||||
```ts
|
||||
import { Recipe } from './schemas';
|
||||
|
||||
const out = await callLLM(prompt);
|
||||
const parsed = Recipe.safeParse(out);
|
||||
expect(parsed.success).toBe(true);
|
||||
if (!parsed.success) console.log(parsed.error);
|
||||
```
|
||||
|
||||
### Latency / cost 추적
|
||||
```ts
|
||||
const start = Date.now();
|
||||
const r = await openai.chat.completions.create({...});
|
||||
const ms = Date.now() - start;
|
||||
const usage = r.usage!;
|
||||
const cost = usage.prompt_tokens * 2.5e-6 + usage.completion_tokens * 1e-5;
|
||||
|
||||
track('llm.eval', { ms, cost, prompt_tokens: usage.prompt_tokens });
|
||||
```
|
||||
|
||||
### CI 회귀
|
||||
```yaml
|
||||
# .github/workflows/llm-eval.yml
|
||||
on: [pull_request]
|
||||
jobs:
|
||||
eval:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- run: npm i
|
||||
- run: npx promptfoo eval --output report.json
|
||||
- run: node scripts/check-regression.js report.json
|
||||
# baseline 점수보다 5% 이상 하락 시 실패
|
||||
```
|
||||
|
||||
## 🤔 의사결정 기준
|
||||
| 출력 종류 | 채점 |
|
||||
|---|---|
|
||||
| Exact answer | exact match / contains |
|
||||
| JSON / 구조 | Schema parse |
|
||||
| 분류 | accuracy / F1 |
|
||||
| 자유 텍스트 | LLM-as-judge / rouge / BLEU |
|
||||
| 비교 (어느 게 나아?) | pairwise A/B |
|
||||
| 실제 사용자 신호 | thumbs up/down / 재질문률 |
|
||||
|
||||
## ❌ 안티패턴
|
||||
- **Eval 없이 prod 배포**: 회귀 검출 불가.
|
||||
- **Test set 작음 (5개)**: 변동 큼. 50+ 권장.
|
||||
- **Test set leak (학습에 사용)**: 거짓 점수.
|
||||
- **LLM-as-judge — 같은 모델로 채점**: 자기 편향.
|
||||
- **Cost / latency 무시**: 정확도만 보면 비용 폭발.
|
||||
- **Production 못 배포 — 매번 eval**: 작은 hot-set 만 매 PR, 큰 건 nightly.
|
||||
- **Subjective only — 자동화 X**: 매번 사람 — 못 scale.
|
||||
|
||||
## 🤖 LLM 활용 힌트
|
||||
- Promptfoo / Braintrust / LangSmith 권장.
|
||||
- LLM-as-judge 는 다른 모델로.
|
||||
- 회귀 5% 임계값 + cost / latency 같이.
|
||||
|
||||
## 🔗 관련 문서
|
||||
- [[AI_Prompt_Engineering_Patterns]]
|
||||
- [[AI_Structured_Output_Zod]]
|
||||
- [[AI_RAG_Pattern_Basics]]
|
||||
Reference in New Issue
Block a user