Files
2nd/10_Wiki/Topic_Programming/Coding/Backend_Health_Check_Patterns.md
T
Antigravity Agent 9148c358d0 docs(10_Wiki): 위키 전체 재구성 — Topic_* 폴더를 4개 카테고리로 통합 + 대규모 중복 제거
Topic_Agent/Topic_Blog/Topics/Topics_Biz/Topics_Meeting/Topics_Rag의 마크다운 지식 문서를
Topic_General/Topic_Programming/Topic_Graphic/Topic_Business 4개 카테고리로 재분류.

- 중복 제거: frontmatter의 status:duplicate/merged + duplicate_of/redirect_to 필드로
  자기 자신을 중복으로 선언한 리다이렉트 stub 1032개 제거, 완전 동일 내용 파일 472개 제거,
  동일 파일명·다른 내용 충돌 시 더 큰(완전한) 버전만 유지(162개 제거) — 총 1639개 중복 제거.
- 분류: 폴더 단위로 명확한 항목(AI_and_ML/Coding/Architecture 등 → Programming,
  Comfyui/Visual_Effects → Graphic, Topics_Biz/Topics_Meeting/사업 등 → Business,
  Poetic_Blog_Writing/창의성/Game_Design 등 → General)은 폴더 우선순위로,
  나머지 혼재 폴더(Topic_Agent/Topic_Blog/Topics 루트/Thinking & Reasoning/Other/UI_UX_Assets)는
  title/tags 키워드 스코어링으로 파일 단위 분류(불명확한 경우 General로 폴백).
  원본 폴더명은 "From_*" 서브폴더로 보존해 추적 가능성 유지.
- 최종 배치: Programming 2784 / General 1608 / Graphic 285 / Business 249 = 4926개 문서.
- 에이전트 운영 상태(.astra/.agent/.obsidian/sessions/memory/_company/docs/lessons/_shared/src)는
  지식 콘텐츠가 아니므로 재분류 대상에서 제외하고 원위치 유지.
- Topics/Topic_email(상위 보호 폴더 Topic_email과 파일명 100% 중복) 삭제 — 보호 폴더 자체는 미변경.
- 완전히 비게 된 Topic_Agent/Topic_Blog/Topics_Biz/Topics_Rag 폴더 제거.
2026-07-05 00:33:48 +09:00

4.0 KiB

id, title, category, status, source_trust_level, verification_status, created_at, updated_at, tags, tech_stack, applied_in, aliases
id title category status source_trust_level verification_status created_at updated_at tags tech_stack applied_in aliases
backend-health-check-patterns Health Check — Liveness vs Readiness Coding draft B conceptual 2026-05-09 2026-05-09
backend
health
kubernetes
observability
vibe-coding
language applicable_to
Any backend / Kubernetes
Backend
liveness probe
readiness probe
startup probe
/healthz

Health Check — Liveness vs Readiness

두 종류 — Liveness = "살아있나? 죽었으면 재시작", Readiness = "트래픽 받을 준비됐나? 안 됐으면 LB에서 빼라". 두 개를 같은 endpoint 로 하면 cascading failure.

📖 핵심 개념

  • Liveness: 프로세스 자체. 단순. 외부 의존성 검사 X.
  • Readiness: 트래픽 처리 가능. DB / 캐시 / 의존 서비스 검사 OK.
  • Startup: 초기화 오래 걸리는 앱용. liveness 무시 기간.

💻 코드 패턴

Express 기준

let isReady = false;

app.get('/healthz', (req, res) => res.status(200).json({ status: 'ok' })); // liveness

app.get('/ready', async (req, res) => {
  if (!isReady) return res.status(503).json({ ready: false });
  const checks = await Promise.allSettled([
    pingDB(),       // 1s timeout
    pingRedis(),
    pingDownstream(),
  ]);
  const failed = checks.filter(c => c.status === 'rejected');
  if (failed.length > 0) {
    return res.status(503).json({ ready: false, failed: failed.length });
  }
  res.status(200).json({ ready: true });
});

// 부팅 끝나면
async function bootstrap() {
  await db.connect();
  await loadCaches();
  await warmup();
  isReady = true;
}

// graceful shutdown
process.on('SIGTERM', async () => {
  isReady = false;          // LB 빼지게
  setTimeout(async () => {  // 진행 중 요청 처리 후 종료
    await db.disconnect();
    process.exit(0);
  }, 30_000);
});

Kubernetes manifest

livenessProbe:
  httpGet: { path: /healthz, port: 8080 }
  periodSeconds: 30
  failureThreshold: 3
  timeoutSeconds: 1
readinessProbe:
  httpGet: { path: /ready, port: 8080 }
  periodSeconds: 5
  failureThreshold: 2
  timeoutSeconds: 3
startupProbe:
  httpGet: { path: /healthz, port: 8080 }
  periodSeconds: 10
  failureThreshold: 30   # 5분 init 허용

Detailed health (옵션)

app.get('/health/detail', async (req, res) => {
  const [db, redis, queue] = await Promise.allSettled([dbCheck(), redisCheck(), queueCheck()]);
  res.status(200).json({
    db: db.status === 'fulfilled' ? db.value : 'down',
    redis: redis.status === 'fulfilled' ? redis.value : 'down',
    queue: queue.status === 'fulfilled' ? queue.value : 'down',
    version: process.env.GIT_SHA,
    uptime: process.uptime(),
  });
});

🤔 의사결정 기준

의존 Readiness 포함
핵심 DB (없으면 절대 안 됨)
보조 서비스 (analytics, search) — degraded 모드
외부 SaaS (이메일 provider) — 큐로 흡수
Cache (Redis) 보통 — fallback to DB
의존 마이크로서비스 depends — 핵심이면

안티패턴

  • Liveness 가 DB ping: DB 일시 장애 → 모든 pod 재시작 → 폭주. liveness 는 프로세스만.
  • Readiness 가 너무 무거움: probe 자체가 부하. 캐시 + 짧은 timeout.
  • graceful shutdown 없음: SIGTERM 즉시 죽음 → 진행 중 요청 cut. preStop hook + drain.
  • probe timeout = period: 마지막 호출이 끝나기 전 다음 호출 → 누적.
  • dependency 트리 따라 cascading liveness: A 가 B 보고, B 가 C 보면 C 일시 장애가 A 까지 재시작.
  • 버전 / git SHA 노출 안 함: 어떤 빌드가 도는지 디버깅 어려움.
  • public / private 구분 안 함: 외부 노출 시 정보 누설. internal /metrics 와 분리.

🤖 LLM 활용 힌트

  • liveness = simple, readiness = dependency-aware 분리.
  • SIGTERM → readiness false → drain → exit 표준 시퀀스.

🔗 관련 문서