v2.2.311~312: 응답 지연 근본 개선(KV 캐시 분리·제2뇌 상주 캐시) + 보고 품질 수술
v2.2.311 — 응답 지연 (실측: 출력 22토큰에 94.7초, 원인은 매 턴 13k+ 토큰 전체 재프리필) - KV 캐시 친화 프롬프트 분리(kvCachePromptSplit, 기본 ON): message[0]을 불변 정적 본문으로 고정, 날짜/RAG/[CONTEXT]/동적 블록을 마지막 user 메시지 직전의 internal system 메시지로 이동 — llama.cpp prompt cache 프리픽스 재사용으로 턴당 재프리필을 "직전 교환 + 동적 컨텍스트"로 축소. truncation 도 tail 적용. - 검색 토큰 예산 현실화(retrievalTokenBudget, 0=자동): 창의 25%(8k~80k) → 12%(2.5k~6k 클램프). - continuation 은 depth-0 memoryCtx 재사용: 라운드당 재검색 3~8초 제거 + 빈 쿼리 재검색으로 청크가 갈리던 문제 제거 + 턴 내 프롬프트 안정화. - 제2뇌 상주 캐시(신규 brainWatch.ts): 재귀 fs.watch 세대 카운터로 변경 없으면 디렉터리 워크·파일별 statSync 전면 생략. 수정 직후 3초 창은 신뢰 제외(이벤트 지연 레이스 가드), 워처 불가 시 종전 폴백, 활성화 시 백그라운드 워밍, 유휴 해제 30분→2시간. v2.2.312 — 보고 품질 (실사례: "## 4."부터 시작하는 7줄 일반론 보고서) - 중간 라운드 본문 표시 버그 수정: 액션과 함께 작성된 섹션(1~3)이 화면에 한 번도 안 나가고 최종 라운드만 표시되던 근본 원인 제거 — stripForDisplay.ts 로 액션 태그만 걷어내고 라운드 순서대로 버블에 표시. - '분석 보고' 업무 유형 신설(requirementGraph): 보고 개요(첫 줄 자기선언)·파일 근거(주장마다 실제 읽은 파일 인용, 일반론 금지)·구조·발견·다음 단계 강제, 유형 감지 시 "최대 3섹션" 규칙보다 필수 요소 커버 우선. - 근거 없는 분석 감지: 파일을 읽고도 인용 2개 미만인 장문 분석에 "근거 인용 없음" footer 경고 (헛조사 감지의 반대 방향). 검증: 전체 테스트 954건 통과(신규 28), tsc 무오류, vsix 패키징·설치 확인. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
+80
-20
@@ -5,6 +5,8 @@ import * as fs from 'fs';
|
||||
import {
|
||||
findBrainFiles,
|
||||
getSystemPrompt,
|
||||
getStaticSystemPrompt,
|
||||
getDateTimeContextBlock,
|
||||
shouldAutoPushBrain,
|
||||
buildApiUrl,
|
||||
getActiveBrainProfile,
|
||||
@@ -322,6 +324,14 @@ export class AgentExecutor {
|
||||
actionStats: { reads: number; lists: number; investigates: number };
|
||||
/** [v2.2.309] 조사 턴 모델 오버라이드 — depth 0 에서 결정, continuation 에도 유지. */
|
||||
investigationModelOverride: string | null;
|
||||
/**
|
||||
* [v2.2.311] depth 0 에서 빌드한 memoryCtx 문자열 캐시 — continuation depth 는
|
||||
* 재검색하지 않고 이걸 재사용한다. 종전엔 depth 마다 buildMemoryContext 를 다시
|
||||
* 돌렸는데, (a) 검색 3~8초가 라운드마다 추가되고 (b) continuation 의 prompt 는
|
||||
* null 이라 *빈 쿼리로* 재검색해 엉뚱한 청크로 갈아끼우고 (c) 프롬프트가 흔들려
|
||||
* KV 캐시도 깨졌다. actionStats 와 같은 이유로 depth 0 진입부에서만 초기화.
|
||||
*/
|
||||
memoryCtxCache: string | null;
|
||||
} = {
|
||||
retrieval: null,
|
||||
lessons: [],
|
||||
@@ -331,6 +341,7 @@ export class AgentExecutor {
|
||||
confidenceSignals: null,
|
||||
actionStats: { reads: 0, lists: 0, investigates: 0 },
|
||||
investigationModelOverride: null,
|
||||
memoryCtxCache: null,
|
||||
};
|
||||
|
||||
/** Per-turn state 일괄 정리. turn 시작/abort/load session 시 호출. */
|
||||
@@ -560,6 +571,7 @@ export class AgentExecutor {
|
||||
// [v2.2.309] turn 전체(모든 depth) 누적 상태 — depth 0 에서만 초기화.
|
||||
this._turnCtx.actionStats = { reads: 0, lists: 0, investigates: 0 };
|
||||
this._turnCtx.investigationModelOverride = null;
|
||||
this._turnCtx.memoryCtxCache = null;
|
||||
}
|
||||
|
||||
// 1. Prepare Context
|
||||
@@ -755,22 +767,29 @@ export class AgentExecutor {
|
||||
? `\n\n${renderSecondBrainTraceContext(secondBrainTrace)}`
|
||||
: '';
|
||||
const retrievalStartMs = Date.now();
|
||||
const memoryCtx = isCasualConversation
|
||||
? ''
|
||||
: await (async () => {
|
||||
this.resetTurnContext();
|
||||
return buildMemoryContextFn({
|
||||
currentPrompt: prompt || '',
|
||||
activeBrain,
|
||||
agentSkillFile: options.agentSkillFile,
|
||||
chatHistory: this.chatHistory,
|
||||
memoryManager: this.memoryManager,
|
||||
retrievalOrchestrator: this.retrievalOrchestrator,
|
||||
context: this.context,
|
||||
currentTaskId: this.currentTaskId,
|
||||
turnCtx: this._turnCtx,
|
||||
});
|
||||
})();
|
||||
// [v2.2.311] continuation depth 는 depth 0 의 memoryCtx 를 재사용 — 재검색
|
||||
// 3~8초 제거 + 빈 쿼리 재검색으로 청크가 갈리는 문제 제거 + 프롬프트 안정화
|
||||
// (dynamicBlocks 등 turnCtx 파생물도 depth 0 것이 그대로 유지된다).
|
||||
let memoryCtx: string;
|
||||
if (isCasualConversation) {
|
||||
memoryCtx = '';
|
||||
} else if (loopDepth > 0 && this._turnCtx.memoryCtxCache !== null) {
|
||||
memoryCtx = this._turnCtx.memoryCtxCache;
|
||||
} else {
|
||||
this.resetTurnContext();
|
||||
memoryCtx = await buildMemoryContextFn({
|
||||
currentPrompt: prompt || '',
|
||||
activeBrain,
|
||||
agentSkillFile: options.agentSkillFile,
|
||||
chatHistory: this.chatHistory,
|
||||
memoryManager: this.memoryManager,
|
||||
retrievalOrchestrator: this.retrievalOrchestrator,
|
||||
context: this.context,
|
||||
currentTaskId: this.currentTaskId,
|
||||
turnCtx: this._turnCtx,
|
||||
});
|
||||
this._turnCtx.memoryCtxCache = memoryCtx;
|
||||
}
|
||||
if (loopDepth === 0 && !isCasualConversation && this._turnCtx.retrieval) {
|
||||
recordTelemetry({
|
||||
kind: 'retrieval',
|
||||
@@ -806,9 +825,19 @@ export class AgentExecutor {
|
||||
? buildPriorTurnConclusionContext(this.chatHistory)
|
||||
: '';
|
||||
// System prompt build (agent vs astra mode) → src/agent/handlePrompt/{buildAgentModeSystemPrompt,buildAstraModeSystemPrompt}.ts
|
||||
const fullSystemPrompt: string = isAgentMode
|
||||
//
|
||||
// [KV 캐시 분리 v2.2.311] 기본 경로(호출자가 systemPrompt 를 넘기지 않은 경우)는
|
||||
// message[0] 을 *정적 본문만* 으로 고정하고, 날짜/RAG/[CONTEXT]/동적 블록 전부를
|
||||
// dynamicContextTail 로 분리해 computeBudgetedRequest 가 마지막 user 메시지 직전에
|
||||
// 삽입한다. llama.cpp prompt cache 가 정적 프롬프트+과거 히스토리를 재사용하게 되어
|
||||
// 매 턴 전체 재프리필(실측 13k 토큰 ≈ 90초)이 "직전 교환 + tail" 로 줄어든다.
|
||||
// 커스텀 systemPrompt 호출자(멀티에이전트 등)는 종전 단일-시스템 경로 유지.
|
||||
const kvSplitEnabled = getConfig().kvCachePromptSplit !== false
|
||||
&& options.systemPrompt === undefined;
|
||||
const builderBasePrompt = kvSplitEnabled ? '' : systemPrompt;
|
||||
const builtSystemPrompt: string = isAgentMode
|
||||
? buildAgentModeSystemPrompt({
|
||||
systemPrompt,
|
||||
systemPrompt: builderBasePrompt,
|
||||
agentSkillContext: options.agentSkillContext || '',
|
||||
modeBridgeCtx,
|
||||
priorConclusionCtx,
|
||||
@@ -824,7 +853,7 @@ export class AgentExecutor {
|
||||
})
|
||||
: buildAstraModeSystemPrompt({
|
||||
prompt,
|
||||
systemPrompt,
|
||||
systemPrompt: builderBasePrompt,
|
||||
modeBridgeCtx,
|
||||
priorConclusionCtx,
|
||||
designerCtx,
|
||||
@@ -839,6 +868,12 @@ export class AgentExecutor {
|
||||
knowledgeMix: this._turnCtx.knowledgeMix,
|
||||
dynamicBlocks: this._turnCtx.dynamicBlocks,
|
||||
});
|
||||
// Split 모드: head = 정적 프롬프트(불변), tail = 날짜 + 빌더 산출(동적 전부).
|
||||
// Legacy 모드: 종전 그대로 head 에 전부.
|
||||
const fullSystemPrompt: string = kvSplitEnabled ? getStaticSystemPrompt() : builtSystemPrompt;
|
||||
const dynamicContextTail: string | undefined = kvSplitEnabled
|
||||
? `${getDateTimeContextBlock()}${builtSystemPrompt}`
|
||||
: undefined;
|
||||
// Context budget computation → src/agent/handlePrompt/computeBudgetedRequest.ts
|
||||
const imageCount = (reqMessages as any[])
|
||||
.reduce((n, m) => n + (Array.isArray(m?.images) ? m.images.length : 0), 0);
|
||||
@@ -874,7 +909,7 @@ export class AgentExecutor {
|
||||
const lastUserIdx = reqMessages.map((m) => m.role).lastIndexOf('user');
|
||||
const lastUser = lastUserIdx >= 0 ? reqMessages[lastUserIdx] : undefined;
|
||||
const content = typeof lastUser?.content === 'string' ? lastUser.content : '';
|
||||
const sysTokens = estimateTokens(fullSystemPrompt) + 4;
|
||||
const sysTokens = estimateTokens(fullSystemPrompt) + (dynamicContextTail ? estimateTokens(dynamicContextTail) : 0) + 4;
|
||||
const mrCfg = {
|
||||
enabled: true,
|
||||
triggerRatio: config.mapReduceTriggerRatio,
|
||||
@@ -934,6 +969,7 @@ export class AgentExecutor {
|
||||
|
||||
const _budget = computeBudgetedRequest({
|
||||
fullSystemPrompt,
|
||||
dynamicContextTail,
|
||||
reqMessages,
|
||||
actualModel,
|
||||
config,
|
||||
@@ -1462,6 +1498,22 @@ export class AgentExecutor {
|
||||
await this.context.workspaceState.update('lastActionStr', currentActionStr);
|
||||
logInfo('Autonomous loop continuing after actions.', { loopDepth: loopDepth + 1, actions: report });
|
||||
|
||||
// [v2.2.312] 중간 라운드 본문 표시 — 액션과 *함께* 작성된 섹션이 화면에서
|
||||
// 증발하던 버그 수정. 종전엔 액션이 있는 라운드는 여기서 return 하며 본문을
|
||||
// 한 번도 webview 에 보내지 않았고(표시는 라이브 스트리밍뿐 — depth 0 전용),
|
||||
// 최종 라운드만 streamChunk 로 붙었다. 그 결과 모델이 히스토리에서 자기 이전
|
||||
// 섹션(1~3)을 보고 "## 4."부터 이어 써서, 사용자에게는 4번부터 시작하는
|
||||
// 보고서가 도착했다 (실사례). 라이브로 이미 표시된 depth 0 는 중복 방지로 제외.
|
||||
if (loopDepth > 0 || !postLiveDeltas) {
|
||||
try {
|
||||
const { stripActionTagsForDisplay } = await import('./agent/actions/stripForDisplay');
|
||||
const roundVisible = stripActionTagsForDisplay(finalAssistantContent);
|
||||
if (roundVisible) {
|
||||
this.webview.postMessage({ type: 'streamChunk', value: `${loopDepth > 0 ? '\n\n' : ''}${roundVisible}` });
|
||||
}
|
||||
} catch { /* 표시 실패가 루프를 막지 않음 */ }
|
||||
}
|
||||
|
||||
// Explicitly tell the AI to look at the results and continue
|
||||
const continuationPrompt = `The requested local action has been executed.\nAction report:\n${report.join('\n')}\nUse the action result messages already in the conversation to answer the user's original request directly, in the user's language. Do not say you are waiting for the next instruction.`;
|
||||
|
||||
@@ -1554,6 +1606,14 @@ export class AgentExecutor {
|
||||
if (hollowInv.hollow) {
|
||||
this.webview.postMessage({ type: 'streamChunk', value: formatHollowInvestigationFooter(hollowInv.fileMentions) });
|
||||
logInfo('Hollow Investigation 감지 (continuation).', { files: hollowInv.fileMentions, stats: this._turnCtx.actionStats });
|
||||
} else {
|
||||
// [v2.2.312] 반대 방향 — 파일을 읽고도 근거 인용 없는 일반론 분석 경고.
|
||||
const { detectUngroundedAnalysis, formatUngroundedAnalysisFooter } = await import('./intelligence/investigationPipeline');
|
||||
const ug = detectUngroundedAnalysis(finalAssistantContent, this._turnCtx.actionStats);
|
||||
if (ug.ungrounded) {
|
||||
this.webview.postMessage({ type: 'streamChunk', value: formatUngroundedAnalysisFooter(ug.readCount) });
|
||||
logInfo('Ungrounded Analysis 감지 (continuation).', { readCount: ug.readCount, fileRefs: ug.fileRefs });
|
||||
}
|
||||
}
|
||||
} catch { /* 감지 실패가 답변을 막지 않음 */ }
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user