Overview
GLRMask is a grammar-constrained generation library for LLM decoding. It moves work ahead of time to keep mask generation as fast as possible.
GLRMask compiles a grammar against a model vocabulary into a Constraint. A compiled Constraint is immutable and reusable. Compilation cost depends on the grammar, optimization choice and machine; the measured JSON Schema distribution appears below.
GLRMask accepts JSON Schema and general context-free grammars in EBNF, Lark, and GLRM. GLRM also supports composition from separately compiled subgrammars, including tokens that cross parent/child boundaries.
Benchmarks
Time between masks
Statistics
Lower is better.
Time between masks
· µs| Statistic | GLRMaskstatic | GLRMaskdynamic | llguidance |
|---|---|---|---|
| mean | 2.9 | 19.0 | 27.5 |
| p50 | 2.5 | 4.4 | 8.7 |
| p90 | 4.5 | 58.8 | 37.9 |
| p99 | 9.5 | 145.0 | 379.6 |
| p99.9 | 22.7 | 525.4 | 1360.9 |
| p99.99 | 36.9 | 2139.0 | 3548.1 |
| max | 59.3 | 3345.2 | 7843.4 |
Time to first mask
· ms| Statistic | GLRMaskstatic | GLRMaskdynamic | llguidance |
|---|---|---|---|
| mean | 27.4 | 2.5 | 0.8 |
| p50 | 9.7 | 1.4 | 0.5 |
| p90 | 62.1 | 4.4 | 1.3 |
| p99 | 229.4 | 21.9 | 6.9 |
| max | 1052.9 | 132.0 | 136.6 |
Measurement details
macOS 27 ARM64, 10 logical CPUs, Rust 1.95.0; Llama 3.1 8B Instruct vocabulary.
TBM uses calling-thread CPU time and the elementwise minimum of two native timing processes. TTFM uses wall time from the first process. Validation is enabled; one problem runs at a time, with a 60-second build timeout. Vocabulary preparation is outside the timer.
The full run contains 9,558 schemas. Static has 8,326 matched schemas; Dynamic and llguidance have 8,327. TBM uses 2,925,983 aligned positions. The schema cohorts differ by one schema.
Full-run outcomes
Incorrect examples are semantic failures; build success does not establish correctness.
| Mode | Build OK | Build errors | Incorrect examples | Runtime-error examples |
|---|---|---|---|---|
| Static | 9325 | 233 | 460 | 0 |
| Dynamic | 9326 | 232 | 460 | 0 |
| llg 1.8.0 | 8332 | 1226 | 100 | 6 |
Time between masks
Statistics
Lower is better.
Time between masks
· µs| Statistic | GLRMaskstatic | GLRMaskdynamic | llguidance |
|---|---|---|---|
| mean | 17.0 | 1322.0 | 1101.9 |
| p50 | 18.8 | 1256.5 | 938.0 |
| p90 | 27.1 | 3193.7 | 2536.2 |
| p99 | 37.6 | 3687.7 | 3125.0 |
| p99.9 | 53.9 | 3941.4 | 3517.2 |
| p99.99 | 66.8 | 4413.6 | 4021.2 |
| max | 67.0 | 4563.1 | 4217.4 |
Time to first mask
· ms| Statistic | GLRMaskstatic | GLRMaskdynamic | llguidance |
|---|---|---|---|
| value | 3674.4 | 94.2 | 5.8 |
Measurement details
macOS ARM64; full JavaScript fixture: 31 examples and 4,099 tokens. GLRMask 5f2609ac; CFA 159759c4; official llguidance 1.8.0. This dataset is separate from JSON Schema.
TBM uses same-pass mask plus commit at the native adapter boundary, measured with calling-thread CPU time, then takes the elementwise minimum. Static has 20 measured traversals; Dynamic and llguidance have one. One warmup traversal is excluded.
The shown first-mask value is the selected grammar-build wall time plus the initial trace’s first-mask CPU time. Build run counts are Static 1, Dynamic 3 and llguidance 3. The 31 derived example values share a grammar build minimum and are retained as evidence, not shown as an independent distribution. The token table uses linear quantiles; the plot shows the inverse empirical distribution.
All 31 examples completed with zero build, mask, commit or target-token failures. Static and Dynamic masks agree. The canonical JavaScript grammar policies differ; cross-framework mask differences are retained in the evidence, and full-language equivalence is not claimed.
Getting started
Installation
python -m pip install glrmask
Usage
GLRMask has three ordinary public layers:
Grammardescribes the language to generate.Constraintis the compiled grammar for one vocabulary, ready to run.ConstraintStateis the mutable state for one generated sequence.
Compile a Grammar into a Constraint, then call constraint.start() to create its ConstraintState.
At runtime, call constraint.start() once per generated sequence. Compute the next-token mask, sample an allowed model token, then commit that token. If the constraint was built with end tokens, those IDs become maskable only when the grammar body is accepting; committing one lets the decoder stop.
state = constraint.start()
end_token_ids = (
configured_end_token_ids
)
while generating:
in parallel:
logits = llm.forward(...)
mask = state.mask()
logits = apply_mask(logits, mask)
token_id = sample(logits)
state.commit_token(
token_id
)
if (
token_id
in end_token_ids
):
break
Python quickstart
python -m pip install \
glrmask \
llama-cpp-python \
torch
import numpy as np
from llama_cpp import Llama
from torch import from_numpy
import torch
import glrmask
llm = Llama(
model_path="model.gguf",
logits_all=True,
)
vocab = (
glrmask.Vocab
.from_llama_cpp(llm)
)
end_token_ids = (
vocab
.llama_cpp_end_token_ids
)
schema = {
"type": "string",
"enum": [
"positive",
"negative",
"neutral",
],
}
grammar = (
glrmask.Grammar
.from_json_schema(schema)
)
constraint = grammar.compile(
vocab,
end_tokens=end_token_ids,
optimization=(
glrmask.Optimization
.AUTO
),
)
prompt = (
"Classify this review: "
"The story "
"dragged badly. "
"Sentiment: "
)
input_tokens = llm.tokenize(
prompt.encode()
)
llm.reset()
llm.eval(input_tokens)
state = constraint.start()
generated = []
for _ in range(64):
logits = llm.scores[
llm.n_tokens - 1
]
mask = state.mask(
llm.n_vocab()
)
logits[~mask] = -np.inf
token_id = (
torch.distributions
.Categorical(
logits=from_numpy(
logits
),
)
.sample()
.item()
)
llm.eval([token_id])
generated.append(token_id)
state.commit_token(
token_id
)
if (
token_id
in end_token_ids
):
break
text = llm.detokenize(
generated
).decode()
print(text)
state.mask() returns a NumPy Boolean array indexed by model token ID. Pass a size when the model’s logits vector is wider than the constraint’s natural token coordinate.
Rust quickstart
use glrmask::{
BuildOptions, Grammar,
Optimization, Vocab,
};
fn main() {
let yes =
b"\"yes\"".to_vec();
let no =
b"\"no\"".to_vec();
let eos = b"<eos>".to_vec();
let entries = vec![
(0, yes),
(1, no),
(2, eos),
];
let vocab =
Vocab::new(entries);
let schema = concat!(
r#"{"type":"# ,
r#""string","# ,
r#""enum":["yes","# ,
r#""no"]}"#,
);
let grammar = Grammar
::from_json_schema(
schema,
);
let constraint = grammar
.compile_with(
&vocab,
BuildOptions
::default()
.end_tokens([2])
.optimization(
Optimization
::Auto,
),
)
.unwrap();
let mut state =
constraint.start();
let mask = state.mask();
state.commit_token(0)
.unwrap();
assert!(
state.is_accepting()
);
}
Rust masks are packed u32 bitsets. Bit token_id % 32 of word token_id / 32 indicates whether that token is allowed.
Bindings
GLRM declares child slots with extern grammar NAME; and exact-token slots with extern token NAME;. Grammar::bind attaches source grammars or vocabulary-qualified exact-token values. Attach compiled children with UnlinkedConstraint::bind:
use glrmask::{
Grammar, Result, Vocab,
};
let parent = Grammar
::from_glrm(
r#"
glrm 1;
start document;
extern grammar payload;
nt document = payload;
"#,
);
let source_child = Grammar
::from_json_schema(
r#"{"type":"null"}"#,
);
let composed = parent
.bind(
"payload",
&source_child,
)?;
let _a = composed
.compile(vocab)?;
Use vocab.token(id) to bind a specific token. The token object records both its ID and its vocabulary, so GLRMask can check that it matches the vocabulary used for compilation.
Grammar, Result, Vocab,
};
let grammar = Grammar
::from_glrm(
r#"
glrm 1;
start message;
extern token TOOL_CALL;
nt message = TOOL_CALL;
"#,
);
let token = vocab
.token(tool_call_id)?;
let grammar = grammar.bind(
"TOOL_CALL", token,
)?;
let _constraint = grammar
.compile(vocab)?;
Bindings are immutable. Calling bind(...) returns a new Grammar or UnlinkedConstraint; the original remains reusable.
Compilation fails if any required bindings are unresolved.
let parent = Grammar
::from_glrm(
r#"
glrm 1;
start document;
extern grammar payload;
nt document = payload;
"#,
);
let result =
parent.compile(vocab);
if let Err(error) = result {
println!("{error}");
}
Compilation error: external grammar "payload" is unbound; compile_unlinked() if a reusable pre-link artifact is intended
Composition is useful when your set of tools changes. You can compile the surrounding grammar and each tool schema once, then reuse them as tools are added or removed. Only the connections between them need to be rebuilt.
Cached parents with UnlinkedConstraint
An UnlinkedConstraint lets you compile the surrounding grammar before filling its slots. A UnlinkedConstraint may remain open, can be saved and loaded, and is deliberately not runnable.
let parent = Grammar
::from_glrm(
r#"
glrm 1;
start document;
extern grammar payload;
nt document = payload;
"#,
);
let host = parent
.compile_unlinked(vocab)?;
let child_a = Grammar
::from_ebnf(
r#"start ::= "a""#,
).compile(vocab)?;
let child_b = Grammar
::from_ebnf(
r#"start ::= "b""#,
).compile(vocab)?;
let a = host.bind(
"payload", &child_a,
)?;
let b = host.bind(
"payload", &child_b,
)?;
let _constraint_a =
a.link_with(
BuildOptions::default()
.optimization(
Optimization
::FastRuntime,
),
)?;
let _constraint_b = b.link()?;
UnlinkedConstraint::bind is compiled-only: it accepts a compiled Constraint or a vocabulary-qualified exact-token value. It does not accept an unresolved UnlinkedConstraint, and it does not parse or compile source children. Link a child first before binding it to the parent. Composition stays deferred until link/link_with, so the final optimization preference can choose the boundary construction strategy.
Python follows the same basic idea.
Choosing an optimization mode
GLRMask provides three options with different trade-offs between compilation speed and mask generation. The best choice depends on how often you reuse a constraint.
FastBuild: prioritize fast compilation. Useful when constraints change frequently or are used for short generations. This is the Dynamic mode in the benchmarks.Balanced: Balanced uses a llguidance-based approach, with GLRMask’s vocabulary analysis reducing how many tokens it needs to consider. It offers a middle ground between the two.FastRuntime: do more work during compilation to generate masks faster. Useful when you reuse the same constraint across many generations. This is the Static mode in the benchmarks.
Auto is the default: let GLRMask choose.
use glrmask::{
BuildOptions,
Optimization,
};
let constraint = grammar
.compile_with(
vocab,
BuildOptions::default()
.optimization(
Optimization
::FastBuild,
),
)?;
Ending generation
Pass the model’s end-token IDs when you compile a constraint. GLRMask allows those tokens only when the generated text satisfies the grammar. Configure them when you want to continue until EOS; stop when the sampled token is a configured end ID, committing that token first.
use glrmask::BuildOptions;
let options =
BuildOptions::default()
.end_tokens([eos_id]);
let constraint = grammar
.compile_with(
vocab, options,
)?;
let mut state =
constraint.start();
loop {
let mask = state.mask();
// Use your model's sampler.
let token_id =
sample(&mask);
state.commit_token(
token_id,
)?;
if token_id == eos_id {
break;
}
}
is_accepting() tells you whether the text generated so far is a complete valid match, which may still permit continuation. To stop at the first complete match, check acceptance before each iteration:
let mut state =
constraint.start();
while !state.is_accepting() {
let mask = state.mask();
let token_id =
sample(&mask);
state.commit_token(
token_id,
)?;
}
An initially accepting grammar stops immediately, without generating a token. is_rejected() means the generated prefix cannot be completed to match the grammar.
When composing constraints, set end_tokens on the final compile or link call; a child constraint’s end tokens do not end the enclosing generation.
Persistence
Both compiled object types are serializable:
use glrmask::{
Constraint,
UnlinkedConstraint,
};
let module_bytes =
host.save();
let host = UnlinkedConstraint
::load(module_bytes)?;
let constraint_bytes =
constraint.save();
let constraint = Constraint
::load(
constraint_bytes
.as_slice(),
)?;
Grammar formats
GLRMask accepts JSON Schema, GLRM, Lark, and EBNF. GLRM is the native composition format. A grammar begins with glrm 1; and a start declaration:
glrm 1;
start value;
t NUMBER = /-?(0|[1-9][0-9]*)/;
nt value = NUMBER | "null";
Declarations use =, epsilon is written as eps, and regexes use full-match semantics. Inline g name = { ... }; grammars and externally bound extern grammar name; slots have the same language semantics, including scope-local ignores.
Special model-token IDs are declared with extern token NAME; and bound with vocabulary-qualified values:
let grammar = Grammar
::from_glrm(
r#"
glrm 1;
start message;
extern token TOOL_CALL;
nt message = TOOL_CALL "lookup()";
"#,
);
let token = vocab.token(
tool_call_token_id,
)?;
let constraint = grammar
.bind("TOOL_CALL", token)?
.compile(vocab)?;
Lark and EBNF also support explicit @token(<id>) atoms when the numeric ID is deliberately part of the grammar source.