Python API

Python API reference.

Extremely fast grammar-constrained decoding for LLMs.

The glrmask package compiles a grammar together with a model vocabulary and exposes an incremental next-token mask for use inside a decoding loop.

Allocator policy

The Python extension uses mimalloc with delayed automatic purging enabled. It does not override MIMALLOC_PURGE_DELAY, whose mimalloc v3 default is 1000 ms.

GLRMask defaults each purge to a memory reset (MADV_FREE on supported Unix systems and MEM_RESET on Windows) rather than a synchronous decommit. Reset pages remain reclaimable by the operating system and reusable by mimalloc, but process RSS may not decrease immediately. This avoids charging immediate page decommit work to an otherwise unrelated runtime allocation.

Most runtime work also uses bounded preallocated parser, tokenizer, accumulator, and mask storage. Ordinary applications therefore need no allocator lifecycle calls or manual trimming. Set MIMALLOC_PURGE_DECOMMITS=1 before importing GLRMask when immediate RSS reduction is more important than allocator tail latency.

The unstable glrmask._internal.mimalloc_purge_delay(), glrmask._internal.mimalloc_purge_decommits(), and glrmask._internal.collect_allocator(force=True) helpers remain available for diagnostics and controlled experiments.

Installation

python -m pip install glrmask

This reference describes the release-preparation source at f9f5c84, with package metadata version 0.1.1. It does not confirm that a matching wheel has been published. Use the source-build instructions below while the release is being prepared.

Quickstart

import glrmask

vocab = glrmask.Vocab.from_dict({
    b"hello": 0,
    b" ": 1,
    b"world": 2,
    b"<eos>": 3,
})
constraint = glrmask.Grammar.from_ebnf(
    'start ::= "hello" " " "world"'
).compile(vocab, end_tokens=[3])
state = constraint.start()

assert state.mask().tolist() == [True, False, False, False]
state.commit_token(0)
assert state.mask().tolist() == [False, True, False, False]
state.commit_token(1)
state.commit_token(2)
assert state.is_accepting()
assert state.mask()[3]
state.commit_token(3)
assert not state.mask().any()

state.mask() returns a NumPy Boolean array indexed by model token ID. Pass state.mask(size) when the model’s logits vector is wider than the constraint’s natural token coordinate.

Public model

The ordinary API has four layers:

  • Grammar: immutable source description with source-grammar/exact-token bindings.
  • UnlinkedConstraint: reusable compiled machinery for one exact vocabulary in pre-link form; it may intentionally remain open.
  • Constraint: closed, rooted, immediately runnable compiled constraint.
  • ConstraintState: mutable state for one generated sequence.

Grammar.bind(...) and UnlinkedConstraint.bind(...) return new values and leave the receiver reusable.

Vocabulary and exact tokens

Create a vocabulary from token bytes to IDs or IDs to bytes:

vocab = glrmask.Vocab.from_dict({b"yes": 0, b"no": 1})
vocab = glrmask.Vocab.from_id_to_bytes({0: b"yes", 1: b"no"})

For llama-cpp-python:

from llama_cpp import Llama

llm = Llama(model_path="model.gguf", logits_all=True)
vocab = glrmask.Vocab.from_llama_cpp(llm)
end_token_ids = vocab.llama_cpp_end_token_ids

from_llama_cpp() keeps EOG, control, unused, and empty-piece IDs as exact-only model tokens even though they are omitted from the byte vocabulary. They can therefore be used with vocab.token(id) / vocab.tokens(ids) without inventing fake bytes, while EOG IDs can be supplied directly as end_tokens.

Use vocab.token(id) or vocab.tokens(ids) for extern token bindings. These values retain the complete vocabulary identity; a binding from an incompatible vocabulary is rejected even if its numeric ID happens to match.

grammar = glrmask.Grammar.from_glrm('''
glrm 1;
start message;
extern token TOOL_CALL;
nt message = TOOL_CALL "lookup()";
''')
grammar = grammar.bind("TOOL_CALL", vocab.token(tool_call_token_id))
constraint = grammar.compile(vocab)

Compile a grammar

Construct descriptions with:

glrmask.Grammar.from_json_schema(schema)
glrmask.Grammar.from_glrm(grammar)
glrmask.Grammar.from_lark(grammar)
glrmask.Grammar.from_ebnf(grammar)

Then compile a complete description:

constraint = grammar.compile(
    vocab,
    end_tokens=end_token_ids,
    optimization=glrmask.Optimization.AUTO,
)

The intent-level optimization choices are:

  • Optimization.AUTO
  • Optimization.FAST_BUILD
  • Optimization.BALANCED
  • Optimization.FAST_RUNTIME

They preserve language semantics and all return the same Constraint type. They do not expose GLRMask’s internal static/dynamic/O1/O2/O3 engines.

Select a table-free parser

parser_backend is independent of optimization. The pinned public Python implementation defaults to ParserBackend.TEMPLATE_DFA. The backend can also be selected explicitly:

constraint = grammar.compile(
    vocab,
    optimization=glrmask.Optimization.FAST_RUNTIME,
    parser_backend=glrmask.ParserBackend.TEMPLATE_DFA,
)
assert constraint.parser_backend == glrmask.ParserBackend.TEMPLATE_DFA

The runtime and saved artifact do not retain an LR table. Built-in grammar compilation may use LR machinery transiently to derive the template relations. The ordinary mask and commit engines remain shared, and the backend selection survives serialization. This is not a promise that every grammar, latency percentile, or load operation is faster.

To supply a parser without an LR grammar, construct glrmask.ParserProgram(definition) from a JSON string or a JSON-serializable mapping describing its acyclic POP/READ/PUSH graphs. The program validates and owns the data; no Python callback executes during token generation. Compile it with a matching ordered list of terminal patterns: bytes means a literal, and str means a regular expression.

program = glrmask.ParserProgram(definition)
constraint = program.compile(
    vocab,
    terminal_patterns,
    optimization=glrmask.Optimization.FAST_RUNTIME,
)

For data-only programs, FAST_RUNTIME requests the shared static mask compiler; FAST_BUILD uses the shared dynamic engine (O1); BALANCED uses the bounded native parser and vocabulary quotient (O2). An oversized static expansion raises an error instead of silently choosing a different mode. See the template parser contract and the executable Python examples for the exact graph format and independent language checks.

Built-in template-backed children with finite embedding transfers support compiled composition. Select parser_backend=glrmask.ParserBackend.TEMPLATE_DFA when linking them. Arbitrary data-only or older artifacts may lack the required embedding transfer; those links raise errors rather than reconstructing an LR table. The current public native API does not support LR-backed construction, loading or runtime access. The validation report includes measured results and the remaining performance and compatibility tradeoffs for that historical implementation; it does not define the current public default.

Bind source children

Source composition stays in the source world:

parent = glrmask.Grammar.from_glrm('''
glrm 1;
start document;
extern grammar payload;
nt document = "{" payload "}";
''')

source_child = glrmask.Grammar.from_json_schema(payload_schema)
a = parent.bind("payload", source_child).compile(vocab)

Grammar.bind does not accept a compiled Constraint. To reuse a compiled parent with request-specific compiled children, use compile_unlinked instead.

Cache an unlinked parent

Use compile_unlinked when a parent is reused across requests:

host = parent.compile_unlinked(vocab)
child = glrmask.Grammar.from_json_schema(payload_schema).compile(vocab)

bound = host.bind("payload", child)
constraint = bound.link(optimization=glrmask.Optimization.FAST_RUNTIME)

UnlinkedConstraint.bind is compiled-only: accepted grammar children are runnable Constraint values (plus ExactToken/ExactTokens for token slots). It does not accept a source Grammar or another UnlinkedConstraint. Composition remains deferred until link, so the terminal optimization preference can choose the link strategy.

Unlinked constraints are serializable:

artifact = host.save()
host = glrmask.UnlinkedConstraint.load(artifact)

Decode

Create one state per generated sequence. With configured end tokens, keep their IDs and stop after committing one:

state = constraint.start()
end_token_ids = (
    configured_end_token_ids
)

while generating:
    mask = state.mask(model_vocab_size)
    token_id = sample_with_mask(logits, mask)
    state.commit_token(token_id)
    if token_id in end_token_ids:
        break

The main state operations are:

  • mask(size=None): return the allowed-token mask.
  • fill_mask(words): fill a caller-owned packed int32/uint32 buffer.
  • commit_token(token_id): advance by one model token.
  • commit_bytes(data): advance by raw bytes.
  • forced(): return a forced token sequence when one can be determined.
  • is_accepting(): the current output is a complete valid match; it may still permit continuation.
  • is_rejected(): irrecoverably invalid prefix.

End tokens

End tokens are final-root policy, not grammar-child semantics:

constraint = grammar.compile(vocab, end_tokens=[eos_id])

An end token is allowed only when the grammar body is accepting. A child’s previous end-token policy is not inherited when that compiled constraint is linked into an unlinked parent.

Constraint persistence

Constraint objects are immutable/shareable and retain their selected backend after loading. Built-in template-backed children also retain their finite embedding transfers for subsequent links with the template backend:

artifact = constraint.save()
constraint = glrmask.Constraint.load(artifact)

Passing vocab= to Constraint.load or UnlinkedConstraint.load is optional and validates/shares an already-existing exact vocabulary object.

For a template-backed Constraint, save_with_external_vocab() omits the model vocabulary. Loading this form requires the original exact vocabulary mapping:

artifact = constraint.save_with_external_vocab()
constraint = glrmask.Constraint.load(artifact, vocab=vocab)

Loading an external-vocabulary artifact without its vocabulary, or with a different mapping, raises an error. Do not treat matching vocabulary size alone as a matching mapping.

Grammar formats

GLRMask accepts JSON Schema, GLRM, Lark, and EBNF. GLRM is the native composition format:

glrm 1;
start value;
t NUMBER = /-?(0|[1-9][0-9]*)/;
nt value = NUMBER | "null";

External compiled children use extern grammar NAME;; exact model-token slots use extern token NAME;. Inline g name = { ... }; grammars and externally bound child grammars have the same language semantics, including scope-local ignores.

Lark and EBNF also support explicit @token(<id>) atoms when a numeric model token ID is deliberately part of the grammar source.

The top-level package intentionally exposes intent-level Optimization, not the historical engine-specific dynamic/vocabulary-partition types. Repository experiments remain available through explicit internal/submodule imports and carry no public compatibility guarantee.

Source builds

From the repository root:

python -m venv .venv
. .venv/bin/activate
python -m pip install ./python

Building from source requires a Rust toolchain and the platform’s native linker and build tools. On Windows, activate the environment with .venv\Scripts\activate.

For the exact public source revision documented here, first obtain the repository at this revision:

git clone https://github.com/IsaacBreen/glrmask.git
cd glrmask
git checkout f9f5c84f4de5d0d0098c5b68b9733f66727dc456

Then run the source-build commands above from that checkout.