docs(10_Wiki): 위키 전체 재구성 — Topic_* 폴더를 4개 카테고리로 통합 + 대규모 중복 제거
Topic_Agent/Topic_Blog/Topics/Topics_Biz/Topics_Meeting/Topics_Rag의 마크다운 지식 문서를 Topic_General/Topic_Programming/Topic_Graphic/Topic_Business 4개 카테고리로 재분류. - 중복 제거: frontmatter의 status:duplicate/merged + duplicate_of/redirect_to 필드로 자기 자신을 중복으로 선언한 리다이렉트 stub 1032개 제거, 완전 동일 내용 파일 472개 제거, 동일 파일명·다른 내용 충돌 시 더 큰(완전한) 버전만 유지(162개 제거) — 총 1639개 중복 제거. - 분류: 폴더 단위로 명확한 항목(AI_and_ML/Coding/Architecture 등 → Programming, Comfyui/Visual_Effects → Graphic, Topics_Biz/Topics_Meeting/사업 등 → Business, Poetic_Blog_Writing/창의성/Game_Design 등 → General)은 폴더 우선순위로, 나머지 혼재 폴더(Topic_Agent/Topic_Blog/Topics 루트/Thinking & Reasoning/Other/UI_UX_Assets)는 title/tags 키워드 스코어링으로 파일 단위 분류(불명확한 경우 General로 폴백). 원본 폴더명은 "From_*" 서브폴더로 보존해 추적 가능성 유지. - 최종 배치: Programming 2784 / General 1608 / Graphic 285 / Business 249 = 4926개 문서. - 에이전트 운영 상태(.astra/.agent/.obsidian/sessions/memory/_company/docs/lessons/_shared/src)는 지식 콘텐츠가 아니므로 재분류 대상에서 제외하고 원위치 유지. - Topics/Topic_email(상위 보호 폴더 Topic_email과 파일명 100% 중복) 삭제 — 보호 폴더 자체는 미변경. - 완전히 비게 된 Topic_Agent/Topic_Blog/Topics_Biz/Topics_Rag 폴더 제거.
This commit is contained in:
@@ -0,0 +1,120 @@
|
||||
---
|
||||
id: observability-red-use-metrics
|
||||
title: RED / USE 메트릭 — 어떤 걸 측정할까
|
||||
category: Coding
|
||||
status: draft
|
||||
source_trust_level: B
|
||||
verification_status: conceptual
|
||||
created_at: 2026-05-09
|
||||
updated_at: 2026-05-09
|
||||
tags: [observability, metrics, sli, slo, vibe-coding]
|
||||
tech_stack: { language: "Prometheus / Grafana", applicable_to: ["Backend"] }
|
||||
applied_in: []
|
||||
aliases: [Rate, Errors, Duration, Utilization, Saturation, four golden signals]
|
||||
---
|
||||
|
||||
# RED / USE 메트릭
|
||||
|
||||
> 측정 안 하면 운영 못 함. **RED (서비스): Rate / Errors / Duration**, **USE (리소스): Utilization / Saturation / Errors**. SRE 가 정의한 표준 출발점.
|
||||
|
||||
## 📖 핵심 개념
|
||||
- **RED** (요청 기반 서비스): 사용자 관점.
|
||||
- **USE** (리소스 기반): CPU/메모리/디스크/네트워크.
|
||||
- **Four Golden Signals** (Google SRE): Latency / Traffic / Errors / Saturation.
|
||||
|
||||
## 💻 코드 패턴
|
||||
|
||||
### prom-client (Node)
|
||||
```ts
|
||||
import client from 'prom-client';
|
||||
client.collectDefaultMetrics(); // CPU/heap/eventloop 자동
|
||||
|
||||
// RED — HTTP
|
||||
const httpReqs = new client.Counter({
|
||||
name: 'http_requests_total',
|
||||
help: 'Total HTTP requests',
|
||||
labelNames: ['method', 'route', 'status'],
|
||||
});
|
||||
|
||||
const httpDur = new client.Histogram({
|
||||
name: 'http_request_duration_seconds',
|
||||
help: 'HTTP request duration',
|
||||
labelNames: ['method', 'route', 'status'],
|
||||
buckets: [0.005, 0.01, 0.05, 0.1, 0.3, 0.5, 1, 2, 5],
|
||||
});
|
||||
|
||||
app.use((req, res, next) => {
|
||||
const end = httpDur.startTimer({ method: req.method, route: req.route?.path ?? 'unknown' });
|
||||
res.on('finish', () => {
|
||||
const status = String(res.statusCode);
|
||||
end({ status });
|
||||
httpReqs.inc({ method: req.method, route: req.route?.path ?? 'unknown', status });
|
||||
});
|
||||
next();
|
||||
});
|
||||
|
||||
app.get('/metrics', async (_, res) => {
|
||||
res.set('Content-Type', client.register.contentType);
|
||||
res.end(await client.register.metrics());
|
||||
});
|
||||
```
|
||||
|
||||
### USE — DB pool
|
||||
```ts
|
||||
const dbPoolUtil = new client.Gauge({
|
||||
name: 'db_pool_utilization',
|
||||
help: 'Active connections / pool size',
|
||||
});
|
||||
const dbPoolSat = new client.Gauge({
|
||||
name: 'db_pool_waiting',
|
||||
help: 'Connections waiting',
|
||||
});
|
||||
|
||||
setInterval(() => {
|
||||
dbPoolUtil.set((pool.totalCount - pool.idleCount) / pool.options.max);
|
||||
dbPoolSat.set(pool.waitingCount);
|
||||
}, 5000);
|
||||
```
|
||||
|
||||
### Latency 분포 — Histogram, not Average
|
||||
```promql
|
||||
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
|
||||
# p95
|
||||
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{route="/api/checkout"}[5m]))
|
||||
```
|
||||
|
||||
### Error rate
|
||||
```promql
|
||||
sum(rate(http_requests_total{status=~"5.."}[5m]))
|
||||
/
|
||||
sum(rate(http_requests_total[5m]))
|
||||
```
|
||||
|
||||
## 🤔 의사결정 기준
|
||||
| 영역 | 메트릭 |
|
||||
|---|---|
|
||||
| HTTP API | RED — 라벨: method, route, status |
|
||||
| 워커 / job queue | RED + queue length / age |
|
||||
| DB / Redis | USE — connections, latency, error rate |
|
||||
| 외부 API 호출 | RED — provider, status |
|
||||
| Business KPI | gauge / counter (signups, paid_orders) |
|
||||
| Infra (Pod CPU) | Kubernetes 자동 |
|
||||
|
||||
## ❌ 안티패턴
|
||||
- **평균만**: 평균은 거짓말. p95 / p99 가 사용자 경험.
|
||||
- **라벨 폭증**: userId / requestId 라벨 → cardinality 폭증 → 메모리 / 비용 폭사. 라벨은 enum 같은 것만.
|
||||
- **histogram bucket 부적합**: 1ms~10s 인데 bucket 이 10s 단위. 의미 없음.
|
||||
- **counter 와 gauge 혼동**: counter 는 monotonic. 감소 안 함.
|
||||
- **리셋 시 0 dump**: counter 는 reset 알아서 처리. gauge 만 직접 set.
|
||||
- **메트릭 / 로그 / trace 따로 봄**: exemplar / trace_id 로 연결.
|
||||
- **알림 임계값 절대값**: 트래픽 변동에 거짓 알림. ratio + window.
|
||||
|
||||
## 🤖 LLM 활용 힌트
|
||||
- 새 endpoint 마다 RED 자동 (middleware).
|
||||
- 외부 의존성마다 별도 RED.
|
||||
- p95/p99 SLO 정의 후 알림.
|
||||
|
||||
## 🔗 관련 문서
|
||||
- [[Observability_Structured_Logging]]
|
||||
- [[Observability_OpenTelemetry]]
|
||||
- [[Backend_Health_Check_Patterns]]
|
||||
Reference in New Issue
Block a user