Getting started

Installation

python -m pip install glrmask

Usage

GLRMask has three ordinary public layers:

  • Grammar describes the language to generate.
  • Constraint is the compiled grammar for one vocabulary, ready to run.
  • ConstraintState is the mutable state for one generated sequence.

Compile a Grammar into a Constraint, then call constraint.start() to create its ConstraintState.

At runtime, call constraint.start() once per generated sequence. Compute the next-token mask, sample an allowed model token, then commit that token. If the constraint was built with end tokens, those IDs become maskable only when the grammar body is accepting; committing one lets the decoder stop.

state = constraint.start()
end_token_ids = (
    configured_end_token_ids
)

while generating:
    in parallel:
        logits = llm.forward(...)
        mask = state.mask()

    logits = apply_mask(logits, mask)
    token_id = sample(logits)
    state.commit_token(
        token_id
    )
    if (
        token_id
        in end_token_ids
    ):
        break

Python quickstart

python -m pip install \
    glrmask \
    llama-cpp-python \
    torch
import numpy as np
from llama_cpp import Llama
from torch import from_numpy
import torch

import glrmask

llm = Llama(
    model_path="model.gguf",
    logits_all=True,
)
vocab = (
    glrmask.Vocab
    .from_llama_cpp(llm)
)
end_token_ids = (
    vocab
    .llama_cpp_end_token_ids
)

schema = {
    "type": "string",
    "enum": [
        "positive",
        "negative",
        "neutral",
    ],
}
grammar = (
    glrmask.Grammar
    .from_json_schema(schema)
)
constraint = grammar.compile(
    vocab,
    end_tokens=end_token_ids,
    optimization=(
        glrmask.Optimization
        .AUTO
    ),
)

prompt = (
    "Classify this review: "
    "The story "
    "dragged badly. "
    "Sentiment: "
)
input_tokens = llm.tokenize(
    prompt.encode()
)
llm.reset()
llm.eval(input_tokens)

state = constraint.start()
generated = []

for _ in range(64):
    logits = llm.scores[
        llm.n_tokens - 1
    ]
    mask = state.mask(
        llm.n_vocab()
    )
    logits[~mask] = -np.inf
    token_id = (
        torch.distributions
        .Categorical(
            logits=from_numpy(
                logits
            ),
        )
        .sample()
        .item()
    )

    llm.eval([token_id])
    generated.append(token_id)
    state.commit_token(
        token_id
    )
    if (
        token_id
        in end_token_ids
    ):
        break

text = llm.detokenize(
    generated
).decode()
print(text)

state.mask() returns a NumPy Boolean array indexed by model token ID. Pass a size when the model’s logits vector is wider than the constraint’s natural token coordinate.

Rust quickstart

use glrmask::{
    BuildOptions, Grammar,
    Optimization, Vocab,
};

fn main() {
    let yes =
        b"\"yes\"".to_vec();
    let no =
        b"\"no\"".to_vec();
    let eos = b"<eos>".to_vec();
    let entries = vec![
        (0, yes),
        (1, no),
        (2, eos),
    ];
    let vocab =
        Vocab::new(entries);
    let schema = concat!(
        r#"{"type":"# ,
        r#""string","# ,
        r#""enum":["yes","# ,
        r#""no"]}"#,
    );
    let grammar = Grammar
        ::from_json_schema(
            schema,
        );
    let constraint = grammar
        .compile_with(
        &vocab,
        BuildOptions
            ::default()
            .end_tokens([2])
            .optimization(
                Optimization
                    ::Auto,
            ),
        )
        .unwrap();

    let mut state =
        constraint.start();
    let mask = state.mask();
    state.commit_token(0)
        .unwrap();
    assert!(
        state.is_accepting()
    );
}

Rust masks are packed u32 bitsets. Bit token_id % 32 of word token_id / 32 indicates whether that token is allowed.

Bindings

GLRM declares child slots with extern grammar NAME; and exact-token slots with extern token NAME;. Grammar::bind attaches source grammars or vocabulary-qualified exact-token values. Attach compiled children with UnlinkedConstraint::bind:

use glrmask::{
    Grammar, Result, Vocab,
};

let parent = Grammar
    ::from_glrm(
    r#"
    glrm 1;
    start document;
    extern grammar payload;
    nt document = payload;
    "#,
);
let source_child = Grammar
    ::from_json_schema(
        r#"{"type":"null"}"#,
    );
let composed = parent
    .bind(
        "payload",
        &source_child,
    )?;

let _a = composed
    .compile(vocab)?;

Use vocab.token(id) to bind a specific token. The token object records both its ID and its vocabulary, so GLRMask can check that it matches the vocabulary used for compilation.

    Grammar, Result, Vocab,
};
let grammar = Grammar
    ::from_glrm(
    r#"
    glrm 1;
    start message;
    extern token TOOL_CALL;
    nt message = TOOL_CALL;
    "#,
);
let token = vocab
    .token(tool_call_id)?;
let grammar = grammar.bind(
    "TOOL_CALL", token,
)?;
let _constraint = grammar
    .compile(vocab)?;

Bindings are immutable. Calling bind(...) returns a new Grammar or UnlinkedConstraint; the original remains reusable.

Compilation fails if any required bindings are unresolved.

let parent = Grammar
    ::from_glrm(
    r#"
    glrm 1;
    start document;
    extern grammar payload;
    nt document = payload;
    "#,
);
let result =
    parent.compile(vocab);
if let Err(error) = result {
    println!("{error}");
}
Compilation error: external grammar "payload" is unbound; compile_unlinked() if a reusable pre-link artifact is intended

Composition is useful when your set of tools changes. You can compile the surrounding grammar and each tool schema once, then reuse them as tools are added or removed. Only the connections between them need to be rebuilt.

Cached parents with UnlinkedConstraint

An UnlinkedConstraint lets you compile the surrounding grammar before filling its slots. A UnlinkedConstraint may remain open, can be saved and loaded, and is deliberately not runnable.

let parent = Grammar
    ::from_glrm(
    r#"
    glrm 1;
    start document;
    extern grammar payload;
    nt document = payload;
    "#,
);
let host = parent
    .compile_unlinked(vocab)?;

let child_a = Grammar
    ::from_ebnf(
        r#"start ::= "a""#,
    ).compile(vocab)?;
let child_b = Grammar
    ::from_ebnf(
        r#"start ::= "b""#,
    ).compile(vocab)?;

let a = host.bind(
    "payload", &child_a,
)?;
let b = host.bind(
    "payload", &child_b,
)?;

let _constraint_a =
    a.link_with(
    BuildOptions::default()
        .optimization(
            Optimization
                ::FastRuntime,
        ),
)?;
let _constraint_b = b.link()?;

UnlinkedConstraint::bind is compiled-only: it accepts a compiled Constraint or a vocabulary-qualified exact-token value. It does not accept an unresolved UnlinkedConstraint, and it does not parse or compile source children. Link a child first before binding it to the parent. Composition stays deferred until link/link_with, so the final optimization preference can choose the boundary construction strategy.

Python follows the same basic idea.

Choosing an optimization mode

GLRMask provides three options with different trade-offs between compilation speed and mask generation. The best choice depends on how often you reuse a constraint.

  • FastBuild: prioritize fast compilation. Useful when constraints change frequently or are used for short generations. This is the Dynamic mode in the benchmarks.
  • Balanced: Balanced uses a llguidance-based approach, with GLRMask’s vocabulary analysis reducing how many tokens it needs to consider. It offers a middle ground between the two.
  • FastRuntime: do more work during compilation to generate masks faster. Useful when you reuse the same constraint across many generations. This is the Static mode in the benchmarks.

Auto is the default: let GLRMask choose.

use glrmask::{
    BuildOptions,
    Optimization,
};

let constraint = grammar
    .compile_with(
    vocab,
    BuildOptions::default()
        .optimization(
            Optimization
                ::FastBuild,
        ),
)?;

Ending generation

Pass the model’s end-token IDs when you compile a constraint. GLRMask allows those tokens only when the generated text satisfies the grammar. Configure them when you want to continue until EOS; stop when the sampled token is a configured end ID, committing that token first.

use glrmask::BuildOptions;

let options =
    BuildOptions::default()
    .end_tokens([eos_id]);
let constraint = grammar
    .compile_with(
        vocab, options,
    )?;
let mut state =
    constraint.start();

loop {
    let mask = state.mask();
    // Use your model's sampler.
    let token_id =
        sample(&mask);
    state.commit_token(
        token_id,
    )?;
    if token_id == eos_id {
        break;
    }
}

is_accepting() tells you whether the text generated so far is a complete valid match, which may still permit continuation. To stop at the first complete match, check acceptance before each iteration:

let mut state =
    constraint.start();

while !state.is_accepting() {
    let mask = state.mask();
    let token_id =
        sample(&mask);
    state.commit_token(
        token_id,
    )?;
}

An initially accepting grammar stops immediately, without generating a token. is_rejected() means the generated prefix cannot be completed to match the grammar.

When composing constraints, set end_tokens on the final compile or link call; a child constraint’s end tokens do not end the enclosing generation.

Persistence

Both compiled object types are serializable:

use glrmask::{
    Constraint,
    UnlinkedConstraint,
};

let module_bytes =
    host.save();
let host = UnlinkedConstraint
    ::load(module_bytes)?;

let constraint_bytes =
    constraint.save();
let constraint = Constraint
    ::load(
        constraint_bytes
            .as_slice(),
    )?;

Grammar formats

GLRMask accepts JSON Schema, GLRM, Lark, and EBNF. GLRM is the native composition format. A grammar begins with glrm 1; and a start declaration:

glrm 1;
start value;

t NUMBER = /-?(0|[1-9][0-9]*)/;
nt value = NUMBER | "null";

Declarations use =, epsilon is written as eps, and regexes use full-match semantics. Inline g name = { ... }; grammars and externally bound extern grammar name; slots have the same language semantics, including scope-local ignores.

Special model-token IDs are declared with extern token NAME; and bound with vocabulary-qualified values:

let grammar = Grammar
    ::from_glrm(
    r#"
    glrm 1;
    start message;
    extern token TOOL_CALL;
    nt message = TOOL_CALL "lookup()";
    "#,
);
let token = vocab.token(
    tool_call_token_id,
)?;
let constraint = grammar
    .bind("TOOL_CALL", token)?
    .compile(vocab)?;

Lark and EBNF also support explicit @token(<id>) atoms when the numeric ID is deliberately part of the grammar source.