Complete llama.cpp C/C++ API reference covering model loading, inference, text generation, embeddings, chat, tokenization, sampling, batching, KV cache, LoRA adapters, and state management...
Comprehensive reference for the llama.cpp C API, documenting all non-deprecated functions and common usage patterns.
llama.cpp is a C/C++ implementation for LLM inference with minimal dependencies and state-of-the-art performance. This skill provides:
See references/workflows.md for complete working examples. Basic workflow:
llama_backend_init() - Initialize backendllama_model_load_from_file() - Load modelllama_init_from_model() - Create contextllama_tokenize() - Convert text to tokensllama_decode() - Process tokensllama_sampler_sample() - Sample next tokenUse this skill when:
llama_model: Loaded model weights and architecturellama_context: Inference state (KV cache, compute buffers)llama_batch: Input tokens and positions for processingllama_batch_ext: Opaque extended batch builder, processed by llama_process() (b11284+)llama_sampler: Token sampling configurationllama_vocab: Vocabulary and tokenizerllama_memory_t: KV cache memory handlellama_backend_init()llama_model_load_from_file()llama_init_from_model()llama_tokenize()llama_encode() or llama_decode()llama_sampler_sample()For detailed API documentation, the complete API is split across 6 files for efficient targeted loading. Start with references/api-core.md which links to all other sections.
API Files:
Total: 219 active functions (b11284) across 6 organized files
Most common: llama_backend_init(), llama_model_load_from_file(), llama_init_from_model(), llama_tokenize(), llama_decode(), llama_sampler_sample(), llama_vocab_is_eog(), llama_memory_clear()
Extended batch (b11284+): llama_batch_ext_init(), llama_batch_ext_add_token(), llama_batch_ext_set_pos(), llama_batch_ext_set_output_logits(), llama_process()
See references/api-core.md for the full API index linking to all function signatures.
See references/workflows.md for 16 complete working examples: basic text generation, chat, embeddings, batch processing, multi-sequence, LoRA, state save/load, custom sampling (XTC/DRY), encoder-decoder models, model detection, memory management patterns, streaming, and the extended batch API (llama_process).
See references/workflows.md for detailed best practices. Key points:
llama_model_default_params(), etc.)llama_n_ctx())llama_vocab_is_eog()End-of-generation check (llama_vocab_is_eog()), logits retrieval (llama_get_logits_ith()), batch creation (llama_batch_get_one()), tokenization buffer handling. See references/workflows.md for complete code examples.
Model loading fails:
n_gpu_layers if GPU memory insufficientTokenization returns negative value:
-n size and retryDecode/encode returns non-zero:
llama_batch_get_one() or llama_batch_init())llama_n_ctx())Silent failures / no output:
llama_vocab_is_eog() immediately returns truellama_log_set()Performance issues:
n_threads for CPUn_gpu_layers for GPU offloadingn_batch for promptsSliding Window Attention (SWA) issues:
ctx_params.swa_full = true to access beyond attention windowllama_model_n_swa(model) to detect SWA size and configuration needsPer-sequence state errors:
llama_state_seq_load_file(ctx, "file", dest_seq_id, ...)Model type detection:
llama_model_has_encoder() before assuming decoder-only architecturellama_encode() then llama_decode() workflowFor advanced issues: https://github.com/ggerganov/llama.cpp/discussions
llama_batch_ext), inference, tokenization, chatllama-cpp.h)b11284 (164 commits since b11120) โ additive, non-breaking: 14 new functions, nothing removed, no signature changes.
struct llama_batch_ext + llama_process(ctx, type, batch) with enum llama_process_type (LLAMA_PROCESS_TYPE_ENCODE/DECODE); return codes match llama_decode(). Builder: llama_batch_ext_init(ctx)/_free/_clear, _add(batch, seq_id)/_add_token/_add_embd (return batch index, or -1 full / -2 invalid token / -3 invalid seq id), _add_seq, _set_embd_token, _set_embd_state (stub), _set_output_embd/_set_output_logits (equivalent for now), _set_pos (caller must set positions; M-RoPE embedding entries take multiple positions). New struct llama_embd { data, n_rows, n_embd }. batch_ext get_logits/get_embeddings are TODO โ read outputs with llama_get_logits_ith(ctx, idx). llama_decode() now converts llama_batch to this internally. See api-inference.md and workflow 16.llama_batch_ext_ptr (+ llama_batch_ext_deleter) in llama-cpp.h.llama_get_causal_attn(ctx) โ getter paired with llama_set_causal_attn().llama_numa_init() marked "TODO: deprecate and make part of llama_backend_init()" (not deprecated yet).Previously (b11120, 252 commits since b10868) โ small, non-breaking C API changes:
llama_adapter_lora_init_from_file_ptr(model, FILE *) โ load a LoRA adapter from an open FILE* at its current position (mirrors llama_model_load_from_file_ptr). See api-advanced.md.llama_model_load_from_file_ptr() reads from the current file position; the GGUF data section is now aligned relative to the GGUF start. For mmap, its absolute offset must be 32-byte aligned.llama_sampler_chain_n() now returns int32_t (was int) โ ABI-identical on all mainstream platforms.LLAMA_VOCAB_TYPE_TEST = 7 (dummy tokenizer for tests).Previously (b10868, 203 commits since b10665) โ no functions added or removed; the only public C API change is a rename:
BREAKING (rename, not behavior):
enum llama_tensor_read_lazy โ enum llama_lazy_mode; its constants LLAMA_TENSOR_READ_LAZY_OFF/AUTO/ON โ LLAMA_LAZY_MODE_OFF/AUTO/ON; and llama_model_params.tensor_read_lazy โ llama_model_params.lazy_mode. Behavior is unchanged โ update call sites setting the old field/enum names. See api-core.md.Previously (b10665, 249 commits since b10416) โ no functions added, removed, or re-signatured. The public C API changed only in struct fields, one new enum, and the state-file versions:
BREAKING (data, not code):
LLAMA_SESSION_VERSION 9 โ 10 and LLAMA_STATE_SEQ_VERSION 2 โ 3 (recurrent-state rollback in ggml_ssm_scan changed the serialized layout). Session/sequence-state files from older builds are rejected โ llama_state_load_file() / llama_state_seq_load_file() fail on them. Regenerate cached sessions; check the return value and fall back to re-ingesting the prompt.New lazy tensor reading (renamed in b10868, see above):
enum llama_tensor_read_lazy (LLAMA_TENSOR_READ_LAZY_OFF/AUTO/ON) + llama_model_params.tensor_read_lazy. Faults in rows of arch-marked tensors on demand instead of reading them whole at load โ cuts resident memory and load latency for models with very large sparsely-used tensors. AUTO applies only to marked tensors > 4 GiB; both AUTO and ON require an mmap load mode. See api-core.md.New quantizer memory cap:
llama_model_quantize_params.max_buf_size (size_t) โ max bytes of tensor rows held in memory at once, 0 = default (8 GiB). Lower it to quantize very large models on memory-constrained machines.Previously (b10416, 158 commits since b10258) โ multi-output backend sampling, versioning, and a DRY sampler signature change:
BREAKING:
llama_sampler_init_dry() no longer takes the int32_t n_ctx_train parameter (previously the 2nd argument, right after vocab). Update all call sites โ see api-sampling.md.New multi-output backend sampling [EXPERIMENTAL]:
llama_context_params.n_outputs_max_per_seq (uint32_t) โ max sampled outputs per sequence in a ubatch (0 = n_outputs_max).llama_sampler_i.backend_init() gained a uint32_t n_outputs_max_per_seq parameter; new backend_reset() / copy_state() callbacks for authors of custom backend samplers.llama_sampler_copy() โ copy mutable sampler state between two same-type/config samplers without losing the destination's bound compute graph.llama_get_sampled_token_ith(): with multiple outputs, sampler state now advances on token acceptance, not on read; accept a contiguous prefix in output order (no gaps).New load-mode default:
LLAMA_LOAD_MODE_AUTO (-1) is the new default for llama_model_params.load_mode (was LLAMA_LOAD_MODE_MMAP). Auto-detects based on device capabilities โ e.g. avoids mmap on iGPUs.New:
llama_version() โ returns the llama.cpp library version string (from CMake's new semantic-versioning setup).Doc correction (no code change): llama_sampler_init_penalties()'s penalty_last_n doc previously said "-1 = context size"; that was never accurate โ negative values are clamped to 0 (disabled). History-based samplers (DRY, penalties) no longer resolve a "full-context window" from training context size, since backend sampling constructs samplers before a llama_context (and its resolved context length) exists.
Not in llama.h/llama-cpp.h, mentioned for awareness: mtmd gained Qwen3-TTS support โ a breaking change to the llama-tts CLI binary (outside this skill's C API scope).
Previously (b10258, 183 commits since b10075):
BREAKING:
llama_sampler_init_penalties() gained a new required first parameter n_vocab (source it via llama_vocab_n_tokens(vocab)). Update all call sites.New load-mode API (replaces three booleans):
enum llama_load_mode (LLAMA_LOAD_MODE_NONE/MMAP/MLOCK/MMAP_MLOCK/DIRECT_IO) + llama_load_mode_name() / llama_load_mode_from_str().llama_model_params.load_mode replaces the removed use_mmap, use_direct_io, and use_mlock boolean fields.llama_model_params.load_mtp (bool) โ whether to load MTP layers.New vocab function:
llama_vocab_get_suppress_tokens() โ model-specific suppress tokens (gguf key tokenizer.ggml.suppress_tokens).Previously (b10075, ~205 commits since b9870): touched the public C API in exactly one place โ new
LLAMA_FTYPE_MOSTLY_Q2_0 = 41 quantization type (CPU backend), PR #24448. Everything else in that range (new
Hy3/hy_v3 model + MTP speculative decoding, server reasoning_budget_tokens, mtmd NUL-truncation fix,
deepseek-ocr v1 multi-tile, /responses streaming timings, etc.) was internal/server-side and did not change
llama.h.
Recent (added in b9859โb9870, PR #25134):
llama_model_ftype() โ returns the model's file type as an enum llama_ftype (e.g. LLAMA_FTYPE_MOSTLY_Q8_0).llama_ftype_name() โ converts an enum llama_ftype to a human-readable string (e.g. "Q8_0", "Q4_K - Medium"). Pair the two to display a loaded model's quantization.Recent (added in b9840, from b9704):
llama_model_n_layer_nextn() โ returns the number of NextN (Multi-Token Prediction / MTP) layers in the model. These speculative next-token prediction layers power MTP-capable architectures such as DeepSeek V3/V4, GLM, Qwen3.5-MoE, and Step3.5. Returns 0 for non-MTP models. The total layer count equals llama_model_n_layer() (effective layers) + this value.Recent (added in b9704) [EXPERIMENTAL]:
llama_context_params.n_outputs_max (uint32_t) โ max outputs in a ubatch (0 = n_batch). Cap it to reserve less output VRAM when you read only a few logits/embeddings per batch (e.g. one output per sequence during generation).llama_context_params.ctx_other (struct llama_context *) โ a source/target/parent context for sharing inference results or llama_memory (KV cache) between two contexts; used by MTP setups such as Gemma4 MTP.llama_set_warmup() deprecated โ perform warmup runs manually instead. It changed graph topology with MoE models (causing extra reallocations) and will be removed in a future release.Still recent (added in b9246) [EXPERIMENTAL]:
enum llama_context_type (LLAMA_CONTEXT_TYPE_DEFAULT/LLAMA_CONTEXT_TYPE_MTP) + llama_context_params.ctx_typellama_context_params.n_rs_seq + llama_n_rs_seq(ctx)LLAMA_STATE_SEQ_FLAGS_NONE (0), LLAMA_STATE_SEQ_FLAGS_ON_DEVICE (2)Stable Since b8809:
llama_model_load_from_file_ptr(), llama_model_init_from_user()LLAMA_SPLIT_MODE_TENSOR (3) for backend-agnostic tensor parallelismllama_sampler_init_adaptive_p()If you're updating old code:
llama_model_load_from_file() instead of llama_load_model_from_file()llama_model_free() instead of llama_free_model()llama_init_from_model() instead of llama_new_context_with_model()llama_vocab_*() functions instead of llama_token_*()llama_state_*() functions instead of deprecated state functionsllama_set_adapters_lora() instead of the removed llama_set_adapter_lora() for LoRA adaptersllama_vocab_bos() instead of llama_vocab_cls() (CLS is equivalent to BOS)llama_sampler_init_grammar_lazy_patterns() instead of llama_sampler_init_grammar_lazy()llama_set_warmup() (deprecated in b9704)See the API reference for complete mappings.