Skip to content

Reuse a per-session context across chat exchanges - #215

Closed
james-333i wants to merge 3 commits into
huggingface:mainfrom
james-333i:feat/llama-session-context
Closed

james-333i wants to merge 3 commits into
huggingface:mainfrom
james-333i:feat/llama-session-context

Conversation

@james-333i

Copy link
Copy Markdown
Contributor

Every generation created a fresh llama_context and prefilled the full conversation from token zero, so multi-turn cost grew with the square of the transcript. This keeps one context alive per session, reuses the longest shared token prefix, and decodes only the remainder. Backends that cannot rewind fall back to a full decode, and any generation error discards the cached context. Stacks on the Gemma 4 PR.

@james-333i
james-333i force-pushed the feat/llama-session-context branch 2 times, most recently from 59f035f to e79f8f1 Compare September 4, 2026 22:56
@mattt
mattt requested a balanced review from Copilot September 8, 2026 16:04

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Decode failures can leave invalid cached state, concurrent diagnostics race, and caching does not persist per session as described.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Adds reusable llama.cpp session contexts to reduce repeated prefill work, alongside stacked Gemma 4 and vision support.

Changes:

  • Reuses cached token prefixes across chat turns.
  • Adds multimodal projector and image generation support.
  • Adds Gemma 4 prompt rendering and integration tests.
File summaries
File Description
LlamaLanguageModel.swift Implements caching, multimodal generation, and Gemma 4 formatting.
LlamaLanguageModelTests.swift Tests context reuse and vision behavior.
LlamaGemma4TemplateTests.swift Tests Gemma 4 prompt rendering.
Review details

Suppressed comments (1)

Sources/AnyLanguageModel/Models/LlamaLanguageModel.swift:1615

  • A multimodal decode failure is treated as normal completion, so callers receive a successful truncated response and streaming finishes without an error. Propagate decodingFailed, consistent with the prompt-evaluation failure above.
                guard llama_decode(context, batch) == 0 else {
                    break
                }
  • Files reviewed: 3/3 changed files
  • Comments generated: 3
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

/// frees a context that is still decoding: it runs on a transient context
/// instead and leaves the cache untouched.
private let sessionContextLock = NSLock()
private var cachedSessionContext: CachedSessionContext?
Comment on lines +1384 to +1385
lastReusedTokenCount = startIndex
lastPrefillTokenCount = promptTokens.count - startIndex
decodedTokens.append(nextToken)
}

recordCachedTokens(decodedTokens, context: context)
@mattt

mattt commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator

Hi @james-333i. Two small things on this one, written up at the end of #213 (review) so you can do the whole stack in one pass: guard the catch eviction on cachedSessionContext?.context == context, and take the lock around the two diagnostic counters. Otherwise this is good to go 🚀

Image segments threw unsupportedFeature because the backend had no
multimodal path, even though the prebuilt llama.cpp binaries ship the
mtmd library and its helpers.

Accept an mmprojPath at initialization and load the projector next to
the model. When a projector is present, prompt formatting replaces
each image segment with the mtmd media marker and collects payloads in
order, then generation tokenizes the marker-annotated prompt with
mtmd_tokenize and evaluates text and image chunks through
mtmd_helper_eval_chunks before sampling continues from the resulting
position. Both respond and streaming support images, and models
without a projector keep rejecting image input.

Adds live tests generating from an embedded test image through both
paths.
Gemma 4's canonical chat template no longer contains the
start_of_turn marker that llama_chat_apply_template keys its Gemma
detection on, so formatting threw encodingFailed for every Gemma 4
GGUF.

When template application fails and the embedded template carries the
Gemma 4 turn syntax, render it directly: turns open with a turn
marker and role, close with the reverse marker, the assistant role is
named model, and generation opens a model turn. The BOS token is
applied during tokenization, and thinking is opt-in in this format so
no suppression is needed.
Every generation created a fresh llama_context and prefilled the full
rendered conversation from token zero, so multi-turn chat cost grew
with the square of the transcript and long conversations spent most of
their time re-decoding history.

Keep one context alive per session for plain chat generations. Each
exchange tokenizes the rendered prompt, keeps the longest token prefix
shared with the context's recorded state, removes diverged state with
llama_memory_seq_rm, and decodes only the remainder. Backends that
cannot rewind, such as recurrent models, fall back to clearing memory
and decoding the full prompt, and appends need no rewind on any
backend. The final prompt token is always re-decoded so sampling has
fresh logits, generated tokens extend the recorded state as they
decode, and any generation error discards the cached context.

Structured generation, image prompts, and encoder models keep
single-use contexts, and clearCachedContext lets consumers free the
cached state under memory pressure.

Adds a live test asserting prefix reuse on the second turn of a
session.
@james-333i
james-333i force-pushed the feat/llama-session-context branch from e79f8f1 to 49a2acf Compare September 9, 2026 16:13
@mattt
mattt requested a balanced review from Copilot September 10, 2026 12:48

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The cache is model-wide rather than per-session, and decode failures can return truncated output while retaining invalid cached state.

Review details

Suppressed comments (3)

Sources/AnyLanguageModel/Models/LlamaLanguageModel.swift:494

  • The PR promises a context per session, but this model-wide singleton retains only one session. With sequential A → B → A requests, each session switch reaches the replacement path and frees the previous session's KV cache, so a shared model loses the advertised multi-turn reuse. Store entries keyed by session identity, with weak-session cleanup and a bounded eviction policy, rather than replacing one global entry.
        private let sessionContextLock = NSLock()
        private var cachedSessionContext: CachedSessionContext?

Sources/AnyLanguageModel/Models/LlamaLanguageModel.swift:1503

  • A nonzero llama_decode result is silently converted into a successful truncated response, after which the context is recorded for reuse. Because no error reaches generateChatText, its error path cannot discard the potentially invalid cache as promised. Throw decodingFailed here so the cached context is invalidated.
                decodedTokens.append(nextToken)
            }

            recordCachedTokens(decodedTokens, context: context)

Sources/AnyLanguageModel/Models/LlamaLanguageModel.swift:1643

  • A failed token decode currently ends multimodal generation as if it completed normally, returning a truncated response instead of the decodingFailed error used for the initial multimodal evaluation. Propagate the failure so callers can distinguish an inference error from normal EOG completion.
                guard llama_decode(context, batch) == 0 else {
                    break
                }
  • Files reviewed: 3/3 changed files
  • Comments generated: 0 new
  • Review effort level: Balanced

@mattt

mattt commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

Thanks for the checkout design, that's the right answer, and the deferred free is handled on every path I could find. Merging now soon! (GitHub couldn't squash this one after the other two, so it's going in as a merge commit)

@mattt

mattt commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

This landed on main via #223, which carries your commit as-is on top of the merged #213 and #214 (GitHub couldn't squash it here after those two went in). Closing this one; thanks again, @james-333i!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants