Getting started
Installation
python -m pip install glrmask
Usage
GLRMask has three ordinary public layers:
Grammardescribes the language to generate.Constraintis the compiled grammar for one vocabulary, ready to run.ConstraintStateis the mutable state for one generated sequence.
Compile a Grammar into a Constraint, then call constraint.start() to create its ConstraintState.
At runtime, call constraint.start() once per generated sequence. Compute the next-token mask, sample an allowed model token, then commit that token. If the constraint was built with end tokens, those IDs become maskable only when the grammar body is accepting; committing one lets the decoder stop.
state = constraint.start()
end_token_ids = (
configured_end_token_ids
)
while generating:
in parallel:
logits = llm.forward(...)
mask = state.mask()
logits = apply_mask(logits, mask)
token_id = sample(logits)
state.commit_token(
token_id
)
if (
token_id
in end_token_ids
):
break
Python quickstart
python -m pip install \
glrmask \
llama-cpp-python \
torch
import numpy as np
from llama_cpp import Llama
from torch import from_numpy
import torch
import glrmask
llm = Llama(
model_path="model.gguf",
logits_all=True,
)
vocab = (
glrmask.Vocab
.from_llama_cpp(llm)
)
end_token_ids = (
vocab
.llama_cpp_end_token_ids
)
schema = {
"type": "string",
"enum": [
"positive",
"negative",
"neutral",
],
}
grammar = (
glrmask.Grammar
.from_json_schema(schema)
)
constraint = grammar.compile(
vocab,
end_tokens=end_token_ids,
optimization=(
glrmask.Optimization
.AUTO
),
)
prompt = (
"Classify this review: "
"The story "
"dragged badly. "
"Sentiment: "
)
input_tokens = llm.tokenize(
prompt.encode()
)
llm.reset()
llm.eval(input_tokens)
state = constraint.start()
generated = []
for _ in range(64):
logits = llm.scores[
llm.n_tokens - 1
]
mask = state.mask(
llm.n_vocab()
)
logits[~mask] = -np.inf
token_id = (
torch.distributions
.Categorical(
logits=from_numpy(
logits
),
)
.sample()
.item()
)
llm.eval([token_id])
generated.append(token_id)
state.commit_token(
token_id
)
if (
token_id
in end_token_ids
):
break
text = llm.detokenize(
generated
).decode()
print(text)
state.mask() returns a NumPy Boolean array indexed by model token ID. Pass a size when the model’s logits vector is wider than the constraint’s natural token coordinate.
Rust quickstart
use glrmask::{
BuildOptions, Grammar,
Optimization, Vocab,
};
fn main() {
let yes =
b"\"yes\"".to_vec();
let no =
b"\"no\"".to_vec();
let eos = b"<eos>".to_vec();
let entries = vec![
(0, yes),
(1, no),
(2, eos),
];
let vocab =
Vocab::new(entries);
let schema = concat!(
r#"{"type":"# ,
r#""string","# ,
r#""enum":["yes","# ,
r#""no"]}"#,
);
let grammar = Grammar
::from_json_schema(
schema,
);
let constraint = grammar
.compile_with(
&vocab,
BuildOptions
::default()
.end_tokens([2])
.optimization(
Optimization
::Auto,
),
)
.unwrap();
let mut state =
constraint.start();
let mask = state.mask();
state.commit_token(0)
.unwrap();
assert!(
state.is_accepting()
);
}
Rust masks are packed u32 bitsets. Bit token_id % 32 of word token_id / 32 indicates whether that token is allowed.
Bindings
GLRM declares child slots with extern grammar NAME; and exact-token slots with extern token NAME;. Grammar::bind attaches source grammars or vocabulary-qualified exact-token values. Attach compiled children with UnlinkedConstraint::bind:
use glrmask::{
Grammar, Result, Vocab,
};
let parent = Grammar
::from_glrm(
r#"
glrm 1;
start document;
extern grammar payload;
nt document = payload;
"#,
);
let source_child = Grammar
::from_json_schema(
r#"{"type":"null"}"#,
);
let composed = parent
.bind(
"payload",
&source_child,
)?;
let _a = composed
.compile(vocab)?;
Use vocab.token(id) to bind a specific token. The token object records both its ID and its vocabulary, so GLRMask can check that it matches the vocabulary used for compilation.
Grammar, Result, Vocab,
};
let grammar = Grammar
::from_glrm(
r#"
glrm 1;
start message;
extern token TOOL_CALL;
nt message = TOOL_CALL;
"#,
);
let token = vocab
.token(tool_call_id)?;
let grammar = grammar.bind(
"TOOL_CALL", token,
)?;
let _constraint = grammar
.compile(vocab)?;
Bindings are immutable. Calling bind(...) returns a new Grammar or UnlinkedConstraint; the original remains reusable.
Compilation fails if any required bindings are unresolved.
let parent = Grammar
::from_glrm(
r#"
glrm 1;
start document;
extern grammar payload;
nt document = payload;
"#,
);
let result =
parent.compile(vocab);
if let Err(error) = result {
println!("{error}");
}
Compilation error: external grammar "payload" is unbound; compile_unlinked() if a reusable pre-link artifact is intended
Composition is useful when your set of tools changes. You can compile the surrounding grammar and each tool schema once, then reuse them as tools are added or removed. Only the connections between them need to be rebuilt.
Cached parents with UnlinkedConstraint
An UnlinkedConstraint lets you compile the surrounding grammar before filling its slots. A UnlinkedConstraint may remain open, can be saved and loaded, and is deliberately not runnable.
let parent = Grammar
::from_glrm(
r#"
glrm 1;
start document;
extern grammar payload;
nt document = payload;
"#,
);
let host = parent
.compile_unlinked(vocab)?;
let child_a = Grammar
::from_ebnf(
r#"start ::= "a""#,
).compile(vocab)?;
let child_b = Grammar
::from_ebnf(
r#"start ::= "b""#,
).compile(vocab)?;
let a = host.bind(
"payload", &child_a,
)?;
let b = host.bind(
"payload", &child_b,
)?;
let _constraint_a =
a.link_with(
BuildOptions::default()
.optimization(
Optimization
::FastRuntime,
),
)?;
let _constraint_b = b.link()?;
UnlinkedConstraint::bind is compiled-only: it accepts a compiled Constraint or a vocabulary-qualified exact-token value. It does not accept an unresolved UnlinkedConstraint, and it does not parse or compile source children. Link a child first before binding it to the parent. Composition stays deferred until link/link_with, so the final optimization preference can choose the boundary construction strategy.
Python follows the same basic idea.
Choosing an optimization mode
GLRMask provides three options with different trade-offs between compilation speed and mask generation. The best choice depends on how often you reuse a constraint.
FastBuild: prioritize fast compilation. Useful when constraints change frequently or are used for short generations. This is the Dynamic mode in the benchmarks.Balanced: Balanced uses a llguidance-based approach, with GLRMask’s vocabulary analysis reducing how many tokens it needs to consider. It offers a middle ground between the two.FastRuntime: do more work during compilation to generate masks faster. Useful when you reuse the same constraint across many generations. This is the Static mode in the benchmarks.
Auto is the default: let GLRMask choose.
use glrmask::{
BuildOptions,
Optimization,
};
let constraint = grammar
.compile_with(
vocab,
BuildOptions::default()
.optimization(
Optimization
::FastBuild,
),
)?;
Ending generation
Pass the model’s end-token IDs when you compile a constraint. GLRMask allows those tokens only when the generated text satisfies the grammar. Configure them when you want to continue until EOS; stop when the sampled token is a configured end ID, committing that token first.
use glrmask::BuildOptions;
let options =
BuildOptions::default()
.end_tokens([eos_id]);
let constraint = grammar
.compile_with(
vocab, options,
)?;
let mut state =
constraint.start();
loop {
let mask = state.mask();
// Use your model's sampler.
let token_id =
sample(&mask);
state.commit_token(
token_id,
)?;
if token_id == eos_id {
break;
}
}
is_accepting() tells you whether the text generated so far is a complete valid match, which may still permit continuation. To stop at the first complete match, check acceptance before each iteration:
let mut state =
constraint.start();
while !state.is_accepting() {
let mask = state.mask();
let token_id =
sample(&mask);
state.commit_token(
token_id,
)?;
}
An initially accepting grammar stops immediately, without generating a token. is_rejected() means the generated prefix cannot be completed to match the grammar.
When composing constraints, set end_tokens on the final compile or link call; a child constraint’s end tokens do not end the enclosing generation.
Persistence
Both compiled object types are serializable:
use glrmask::{
Constraint,
UnlinkedConstraint,
};
let module_bytes =
host.save();
let host = UnlinkedConstraint
::load(module_bytes)?;
let constraint_bytes =
constraint.save();
let constraint = Constraint
::load(
constraint_bytes
.as_slice(),
)?;
Grammar formats
GLRMask accepts JSON Schema, GLRM, Lark, and EBNF. GLRM is the native composition format. A grammar begins with glrm 1; and a start declaration:
glrm 1;
start value;
t NUMBER = /-?(0|[1-9][0-9]*)/;
nt value = NUMBER | "null";
Declarations use =, epsilon is written as eps, and regexes use full-match semantics. Inline g name = { ... }; grammars and externally bound extern grammar name; slots have the same language semantics, including scope-local ignores.
Special model-token IDs are declared with extern token NAME; and bound with vocabulary-qualified values:
let grammar = Grammar
::from_glrm(
r#"
glrm 1;
start message;
extern token TOOL_CALL;
nt message = TOOL_CALL "lookup()";
"#,
);
let token = vocab.token(
tool_call_token_id,
)?;
let constraint = grammar
.bind("TOOL_CALL", token)?
.compile(vocab)?;
Lark and EBNF also support explicit @token(<id>) atoms when the numeric ID is deliberately part of the grammar source.