v2.2.311~312: 응답 지연 근본 개선(KV 캐시 분리·제2뇌 상주 캐시) + 보고 품질 수술

v2.2.311 — 응답 지연 (실측: 출력 22토큰에 94.7초, 원인은 매 턴 13k+ 토큰 전체 재프리필)
- KV 캐시 친화 프롬프트 분리(kvCachePromptSplit, 기본 ON): message[0]을 불변
  정적 본문으로 고정, 날짜/RAG/[CONTEXT]/동적 블록을 마지막 user 메시지 직전의
  internal system 메시지로 이동 — llama.cpp prompt cache 프리픽스 재사용으로
  턴당 재프리필을 "직전 교환 + 동적 컨텍스트"로 축소. truncation 도 tail 적용.
- 검색 토큰 예산 현실화(retrievalTokenBudget, 0=자동): 창의 25%(8k~80k) →
  12%(2.5k~6k 클램프).
- continuation 은 depth-0 memoryCtx 재사용: 라운드당 재검색 3~8초 제거 +
  빈 쿼리 재검색으로 청크가 갈리던 문제 제거 + 턴 내 프롬프트 안정화.
- 제2뇌 상주 캐시(신규 brainWatch.ts): 재귀 fs.watch 세대 카운터로 변경 없으면
  디렉터리 워크·파일별 statSync 전면 생략. 수정 직후 3초 창은 신뢰 제외(이벤트
  지연 레이스 가드), 워처 불가 시 종전 폴백, 활성화 시 백그라운드 워밍,
  유휴 해제 30분→2시간.

v2.2.312 — 보고 품질 (실사례: "## 4."부터 시작하는 7줄 일반론 보고서)
- 중간 라운드 본문 표시 버그 수정: 액션과 함께 작성된 섹션(1~3)이 화면에 한 번도
  안 나가고 최종 라운드만 표시되던 근본 원인 제거 — stripForDisplay.ts 로 액션
  태그만 걷어내고 라운드 순서대로 버블에 표시.
- '분석 보고' 업무 유형 신설(requirementGraph): 보고 개요(첫 줄 자기선언)·파일
  근거(주장마다 실제 읽은 파일 인용, 일반론 금지)·구조·발견·다음 단계 강제,
  유형 감지 시 "최대 3섹션" 규칙보다 필수 요소 커버 우선.
- 근거 없는 분석 감지: 파일을 읽고도 인용 2개 미만인 장문 분석에 "근거 인용
  없음" footer 경고 (헛조사 감지의 반대 방향).

검증: 전체 테스트 954건 통과(신규 28), tsc 무오류, vsix 패키징·설치 확인.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-20 13:21:41 +09:00
co-authored by Claude Fable 5
parent a31d273bfe
commit 7d6b8b509f
20 changed files with 834 additions and 38 deletions
+80 -20
View File
@@ -5,6 +5,8 @@ import * as fs from 'fs';
import {
findBrainFiles,
getSystemPrompt,
getStaticSystemPrompt,
getDateTimeContextBlock,
shouldAutoPushBrain,
buildApiUrl,
getActiveBrainProfile,
@@ -322,6 +324,14 @@ export class AgentExecutor {
actionStats: { reads: number; lists: number; investigates: number };
/** [v2.2.309] 조사 턴 모델 오버라이드 — depth 0 에서 결정, continuation 에도 유지. */
investigationModelOverride: string | null;
/**
* [v2.2.311] depth 0 에서 빌드한 memoryCtx 문자열 캐시 — continuation depth 는
* 재검색하지 않고 이걸 재사용한다. 종전엔 depth 마다 buildMemoryContext 를 다시
* 돌렸는데, (a) 검색 3~8초가 라운드마다 추가되고 (b) continuation 의 prompt 는
* null 이라 *빈 쿼리로* 재검색해 엉뚱한 청크로 갈아끼우고 (c) 프롬프트가 흔들려
* KV 캐시도 깨졌다. actionStats 와 같은 이유로 depth 0 진입부에서만 초기화.
*/
memoryCtxCache: string | null;
} = {
retrieval: null,
lessons: [],
@@ -331,6 +341,7 @@ export class AgentExecutor {
confidenceSignals: null,
actionStats: { reads: 0, lists: 0, investigates: 0 },
investigationModelOverride: null,
memoryCtxCache: null,
};
/** Per-turn state 일괄 정리. turn 시작/abort/load session 시 호출. */
@@ -560,6 +571,7 @@ export class AgentExecutor {
// [v2.2.309] turn 전체(모든 depth) 누적 상태 — depth 0 에서만 초기화.
this._turnCtx.actionStats = { reads: 0, lists: 0, investigates: 0 };
this._turnCtx.investigationModelOverride = null;
this._turnCtx.memoryCtxCache = null;
}
// 1. Prepare Context
@@ -755,22 +767,29 @@ export class AgentExecutor {
? `\n\n${renderSecondBrainTraceContext(secondBrainTrace)}`
: '';
const retrievalStartMs = Date.now();
const memoryCtx = isCasualConversation
? ''
: await (async () => {
this.resetTurnContext();
return buildMemoryContextFn({
currentPrompt: prompt || '',
activeBrain,
agentSkillFile: options.agentSkillFile,
chatHistory: this.chatHistory,
memoryManager: this.memoryManager,
retrievalOrchestrator: this.retrievalOrchestrator,
context: this.context,
currentTaskId: this.currentTaskId,
turnCtx: this._turnCtx,
});
})();
// [v2.2.311] continuation depth 는 depth 0 의 memoryCtx 를 재사용 — 재검색
// 3~8초 제거 + 빈 쿼리 재검색으로 청크가 갈리는 문제 제거 + 프롬프트 안정화
// (dynamicBlocks 등 turnCtx 파생물도 depth 0 것이 그대로 유지된다).
let memoryCtx: string;
if (isCasualConversation) {
memoryCtx = '';
} else if (loopDepth > 0 && this._turnCtx.memoryCtxCache !== null) {
memoryCtx = this._turnCtx.memoryCtxCache;
} else {
this.resetTurnContext();
memoryCtx = await buildMemoryContextFn({
currentPrompt: prompt || '',
activeBrain,
agentSkillFile: options.agentSkillFile,
chatHistory: this.chatHistory,
memoryManager: this.memoryManager,
retrievalOrchestrator: this.retrievalOrchestrator,
context: this.context,
currentTaskId: this.currentTaskId,
turnCtx: this._turnCtx,
});
this._turnCtx.memoryCtxCache = memoryCtx;
}
if (loopDepth === 0 && !isCasualConversation && this._turnCtx.retrieval) {
recordTelemetry({
kind: 'retrieval',
@@ -806,9 +825,19 @@ export class AgentExecutor {
? buildPriorTurnConclusionContext(this.chatHistory)
: '';
// System prompt build (agent vs astra mode) → src/agent/handlePrompt/{buildAgentModeSystemPrompt,buildAstraModeSystemPrompt}.ts
const fullSystemPrompt: string = isAgentMode
//
// [KV 캐시 분리 v2.2.311] 기본 경로(호출자가 systemPrompt 를 넘기지 않은 경우)는
// message[0] 을 *정적 본문만* 으로 고정하고, 날짜/RAG/[CONTEXT]/동적 블록 전부를
// dynamicContextTail 로 분리해 computeBudgetedRequest 가 마지막 user 메시지 직전에
// 삽입한다. llama.cpp prompt cache 가 정적 프롬프트+과거 히스토리를 재사용하게 되어
// 매 턴 전체 재프리필(실측 13k 토큰 ≈ 90초)이 "직전 교환 + tail" 로 줄어든다.
// 커스텀 systemPrompt 호출자(멀티에이전트 등)는 종전 단일-시스템 경로 유지.
const kvSplitEnabled = getConfig().kvCachePromptSplit !== false
&& options.systemPrompt === undefined;
const builderBasePrompt = kvSplitEnabled ? '' : systemPrompt;
const builtSystemPrompt: string = isAgentMode
? buildAgentModeSystemPrompt({
systemPrompt,
systemPrompt: builderBasePrompt,
agentSkillContext: options.agentSkillContext || '',
modeBridgeCtx,
priorConclusionCtx,
@@ -824,7 +853,7 @@ export class AgentExecutor {
})
: buildAstraModeSystemPrompt({
prompt,
systemPrompt,
systemPrompt: builderBasePrompt,
modeBridgeCtx,
priorConclusionCtx,
designerCtx,
@@ -839,6 +868,12 @@ export class AgentExecutor {
knowledgeMix: this._turnCtx.knowledgeMix,
dynamicBlocks: this._turnCtx.dynamicBlocks,
});
// Split 모드: head = 정적 프롬프트(불변), tail = 날짜 + 빌더 산출(동적 전부).
// Legacy 모드: 종전 그대로 head 에 전부.
const fullSystemPrompt: string = kvSplitEnabled ? getStaticSystemPrompt() : builtSystemPrompt;
const dynamicContextTail: string | undefined = kvSplitEnabled
? `${getDateTimeContextBlock()}${builtSystemPrompt}`
: undefined;
// Context budget computation → src/agent/handlePrompt/computeBudgetedRequest.ts
const imageCount = (reqMessages as any[])
.reduce((n, m) => n + (Array.isArray(m?.images) ? m.images.length : 0), 0);
@@ -874,7 +909,7 @@ export class AgentExecutor {
const lastUserIdx = reqMessages.map((m) => m.role).lastIndexOf('user');
const lastUser = lastUserIdx >= 0 ? reqMessages[lastUserIdx] : undefined;
const content = typeof lastUser?.content === 'string' ? lastUser.content : '';
const sysTokens = estimateTokens(fullSystemPrompt) + 4;
const sysTokens = estimateTokens(fullSystemPrompt) + (dynamicContextTail ? estimateTokens(dynamicContextTail) : 0) + 4;
const mrCfg = {
enabled: true,
triggerRatio: config.mapReduceTriggerRatio,
@@ -934,6 +969,7 @@ export class AgentExecutor {
const _budget = computeBudgetedRequest({
fullSystemPrompt,
dynamicContextTail,
reqMessages,
actualModel,
config,
@@ -1462,6 +1498,22 @@ export class AgentExecutor {
await this.context.workspaceState.update('lastActionStr', currentActionStr);
logInfo('Autonomous loop continuing after actions.', { loopDepth: loopDepth + 1, actions: report });
// [v2.2.312] 중간 라운드 본문 표시 — 액션과 *함께* 작성된 섹션이 화면에서
// 증발하던 버그 수정. 종전엔 액션이 있는 라운드는 여기서 return 하며 본문을
// 한 번도 webview 에 보내지 않았고(표시는 라이브 스트리밍뿐 — depth 0 전용),
// 최종 라운드만 streamChunk 로 붙었다. 그 결과 모델이 히스토리에서 자기 이전
// 섹션(1~3)을 보고 "## 4."부터 이어 써서, 사용자에게는 4번부터 시작하는
// 보고서가 도착했다 (실사례). 라이브로 이미 표시된 depth 0 는 중복 방지로 제외.
if (loopDepth > 0 || !postLiveDeltas) {
try {
const { stripActionTagsForDisplay } = await import('./agent/actions/stripForDisplay');
const roundVisible = stripActionTagsForDisplay(finalAssistantContent);
if (roundVisible) {
this.webview.postMessage({ type: 'streamChunk', value: `${loopDepth > 0 ? '\n\n' : ''}${roundVisible}` });
}
} catch { /* 표시 실패가 루프를 막지 않음 */ }
}
// Explicitly tell the AI to look at the results and continue
const continuationPrompt = `The requested local action has been executed.\nAction report:\n${report.join('\n')}\nUse the action result messages already in the conversation to answer the user's original request directly, in the user's language. Do not say you are waiting for the next instruction.`;
@@ -1554,6 +1606,14 @@ export class AgentExecutor {
if (hollowInv.hollow) {
this.webview.postMessage({ type: 'streamChunk', value: formatHollowInvestigationFooter(hollowInv.fileMentions) });
logInfo('Hollow Investigation 감지 (continuation).', { files: hollowInv.fileMentions, stats: this._turnCtx.actionStats });
} else {
// [v2.2.312] 반대 방향 — 파일을 읽고도 근거 인용 없는 일반론 분석 경고.
const { detectUngroundedAnalysis, formatUngroundedAnalysisFooter } = await import('./intelligence/investigationPipeline');
const ug = detectUngroundedAnalysis(finalAssistantContent, this._turnCtx.actionStats);
if (ug.ungrounded) {
this.webview.postMessage({ type: 'streamChunk', value: formatUngroundedAnalysisFooter(ug.readCount) });
logInfo('Ungrounded Analysis 감지 (continuation).', { readCount: ug.readCount, fileRefs: ug.fileRefs });
}
}
} catch { /* 감지 실패가 답변을 막지 않음 */ }
}