Backends¶
llenvs supports multiple inference backends with a unified interface.
Lifecycle¶
Backends are closeable and may be used as context managers:
from llenvs.inference.backends import OpenAIBackend
with OpenAIBackend(model="gpt-4o") as backend:
result = backend.generate_chat(...)
Callers own backend lifecycle. Local backends release heavyweight model
resources best-effort on close(), and API backends close reusable client
sessions. Repeated close() calls are safe.
OpenAI, OpenRouter and Anthropic batch calls reuse a dedicated background event
loop for the lifetime of each backend's async client. Chat and tool batches share
that loop, so pooled HTTP connections remain usable across successive calls,
including calls from different threads or from an existing async context. These
public methods are synchronous: they block their caller until the batch finishes.
Concurrent callers on one backend are serialized; requests within each batch use
max_concurrency and results retain input order. close() closes the async client
on its owning loop and stops the loop thread. Always close backends when finished.
Available Backends¶
| Backend | Package | Features |
|---|---|---|
| vLLM | vllm |
Logprobs, batching, prefix continuation, streaming, vision (auto-detected VLMs) |
| HuggingFace | transformers |
Logprobs, batching, prefix continuation, chat templates, vision (auto-detected VLMs) |
| OpenAI | openai |
Logprobs, batching (concurrent), streaming, function calling, vision |
| Anthropic | anthropic |
Batching (concurrent), prefix continuation (prefill), streaming, vision |
| OpenRouter | openai |
Batching (concurrent), access to multiple models, vision |
| LiteLLM | litellm |
Batching (concurrent), 100+ providers via one SDK, function calling, vision, logprobs (provider-dependent), thinking budgets |
| Codex CLI | codex |
Chat, batching (subprocess concurrency), isolated temp workspace, read-only sandbox |
| Claude Code CLI | claude (Claude Code v2.1+) |
Chat, batching (subprocess concurrency), isolated temp workspace, tools disabled, reasoning effort |
vLLM (Local Inference)¶
vLLM runs in-process — no external server needed. Just point it at a model and a GPU.
YAML Config¶
model:
backend: vllm
model: meta-llama/Llama-3.1-8B-Instruct
params:
tensor_parallel_size: 2
dtype: bfloat16
max_model_len: 4096
gpu_memory_utilization: 0.9
Python API¶
from llenvs.inference.backends import VLLMBackend
backend = VLLMBackend(
model_path="meta-llama/Llama-3.1-8B-Instruct",
tensor_parallel_size=2, # Use 2 GPUs
dtype="bfloat16",
max_model_len=4096,
gpu_memory_utilization=0.9,
)
# Batch generation
results = backend.generate(
["Question 1?", "Question 2?", "Question 3?"],
SamplingParams(temperature=0.0),
)
# With logprobs
results = backend.generate_with_logprobs(
["What is 2+2?"],
SamplingParams(temperature=0.0),
num_logprobs=5,
)
for token_lp in results[0].token_logprobs:
print(f"{token_lp.token}: {token_lp.logprob:.3f}")
Capabilities¶
| Feature | Supported | Notes |
|---|---|---|
| logprobs | ✓ | Native support |
| prefix_continuation | ✓ | Native support |
| batching | ✓ | Single batched GPU call |
| streaming | ✓ | Via async iterator |
| chat | ✓ | Via tokenizer.apply_chat_template() |
| function_calling | ✗ | Not supported |
| vision | Auto | VLMs detected automatically |
Tip: If you already have a remote
vllm serveinstance running, useOpenAIBackend(model="your-model", base_url="http://host:port/v1")instead — vLLM's server exposes an OpenAI-compatible API.
Container-hosted continuation scoring¶
SingularityVLLMBackend (vllm_singularity) uses the container server's
/v1/completions endpoint for score_chat / score_chat_batch, with
prompt_logprobs and thinking disabled when rendering the scoring prompt.
The host loads only the tokenizer; model inference runs inside the container.
If any concurrent scoring request fails, score_chat_batch raises
PartialBatchError. Its results retains successful ScoringResult entries and
exceptions in input order; failures maps original input indices to exceptions.
Empty continuations receive empty scoring results without a server request and
do not shift failure indices. Callers can retry failed entries without discarding
successful siblings. The single-request score_chat path uses the same contract.
HuggingFace Transformers (Local Inference)¶
A lightweight alternative to vLLM for running HuggingFace models directly with the transformers library.
from llenvs.inference import SamplingParams
from llenvs.inference.backends import HuggingFaceBackend
backend = HuggingFaceBackend(
model_path="gpt2", # or any HuggingFace model
device="auto", # "cuda", "mps", "cpu", or "auto"
dtype="auto", # "float16", "bfloat16", "float32", or "auto"
)
# Basic generation
results = backend.generate(
["Hello, world!"],
SamplingParams(max_tokens=50, temperature=0.7),
)
print(results[0].text)
# With logprobs
results = backend.generate_with_logprobs(
["The capital of France is"],
SamplingParams(temperature=0.0),
num_logprobs=5,
)
for token_lp in results[0].token_logprobs:
print(f"{token_lp.token}: {token_lp.logprob:.3f}")
Advanced Options¶
# Multi-GPU with device_map
backend = HuggingFaceBackend(
model_path="meta-llama/Llama-3.1-8B-Instruct",
device_map="auto", # Automatically distribute across GPUs
dtype="bfloat16",
)
# With torch.compile optimization
backend = HuggingFaceBackend(
model_path="gpt2",
torch_compile=True, # Enable torch.compile for faster inference
)
# Pass additional kwargs to from_pretrained
backend = HuggingFaceBackend(
model_path="gpt2",
trust_remote_code=True,
revision="main",
)
Chat Generation¶
from llenvs.inference import ChatMessage
# Uses tokenizer.apply_chat_template() internally
result = backend.generate_chat(
[
ChatMessage(role="system", content="You are helpful."),
ChatMessage(role="user", content="Hello!"),
],
SamplingParams(temperature=0.7),
)
Prefix Continuation¶
# Generate multiple continuations from a prefix
continuations = backend.continue_from_prefix(
prefix="Once upon a time",
params=SamplingParams(temperature=0.8, max_tokens=100),
num_continuations=3,
)
for i, cont in enumerate(continuations):
print(f"Continuation {i + 1}: {cont.text}")
Capabilities¶
| Feature | Supported | Notes |
|---|---|---|
| logprobs | ✓ | Via output_scores=True |
| prefix_continuation | ✓ | Native support |
| batching | ✓ | With attention masks |
| streaming | ✗ | Would require TextIteratorStreamer |
| chat | ✓ | Via tokenizer.apply_chat_template() |
| function_calling | ✗ | Not directly supported |
| vision | Auto | VLMs detected and loaded with AutoProcessor + AutoModelForVision2Seq |
OpenAI¶
from llenvs.inference.backends import OpenAIBackend
# Uses OPENAI_API_KEY env var by default
backend = OpenAIBackend(model="gpt-4o")
# Or explicit key with concurrency control
backend = OpenAIBackend(
model="gpt-4o",
api_key="sk-...",
organization="org-...",
max_concurrency=32, # Max concurrent API requests (default: 64)
)
# Chat generation
from llenvs.inference import ChatMessage
result = backend.generate_chat(
[
ChatMessage(role="system", content="You are helpful."),
ChatMessage(role="user", content="Hello!"),
],
SamplingParams(temperature=0.7),
)
Anthropic¶
from llenvs.inference.backends import AnthropicBackend
backend = AnthropicBackend(
model="claude-sonnet-4-20250514",
max_concurrency=32, # Max concurrent API requests (default: 64)
)
# Supports prefix continuation via assistant prefill
continuations = backend.continue_from_prefix(
prefix="Let me solve this step by step:\n1.",
params=SamplingParams(temperature=0.5),
num_continuations=3,
)
OpenRouter¶
from llenvs.inference.backends import OpenRouterBackend
backend = OpenRouterBackend(
model="anthropic/claude-sonnet-4-20250514",
site_url="https://mysite.com",
app_name="MyApp",
max_concurrency=32, # Max concurrent API requests (default: 64)
)
Reasoning / Thinking Budget¶
SamplingParams.thinking_budget is forwarded as
extra_body.reasoning.max_tokens on the chat-completion request, which
OpenRouter routes to the upstream provider's native budget knob
(Anthropic thinking.budget_tokens, Gemini thinking budget, Qwen
thinking_budget); for effort-only models (OpenAI o-series, Grok)
OpenRouter derives an effort bucket from the token count. Non-reasoning
models silently ignore the field. The other thinking-budget knobs
(thinking_budget_per_block, thinking_budget_suffix,
thinking_budget_soft_ratio) require token-level logits intervention
and are inert on this backend — they only take effect for vllm and
huggingface.
To override the derived reasoning block (e.g. send effort instead
of max_tokens, or add exclude: true), pass
extra={"extra_body": {"reasoning": {...}}} on SamplingParams. The
caller-supplied dict shallow-merges over the derived one and wins on
key collisions, so it is also the escape hatch for
reasoning.effort/reasoning.exclude/reasoning.enabled.
OpenRouter returns hidden reasoning separately from visible assistant
content. GenerationResult.text contains only message.content, while
diagnostics are preserved in GenerationResult.metadata: normalized
finish_reason, raw native_finish_reason when present,
completion_tokens_details, top-level reasoning_tokens, and lightweight
presence/count/length fields for message.reasoning and
message.reasoning_details. The full reasoning payload is not copied into
metadata.
OpenRouter responses with a top-level or per-choice error, or a raw/native
finish_reason of error, raise an exception even if HTTP succeeded
and choices contain plausible text or tool calls. No partial answer is returned
as a successful generation. Unnormalized failures raise MalformedResponseError,
which retains provider_error, model/backend
identifiers and a numeric status_code when the error payload provides one.
Callers should distinguish transient failures from permanent 4xx errors rather
than retry every malformed response blindly. Chat rate-limit handling recognizes
429; permanent 4xx codes do not enter its rate-limit wait loop. Concurrent batches
preserve successful siblings and original failed indices in PartialBatchError.
HTTP and body-level 400 errors share input normalization across chat/tool and
single/batch paths: recognized context-limit errors become PromptTooLongError,
and invalid media becomes RecoverableInputError. Authentication, billing,
permission and policy failures are not converted into skippable input failures.
Normalized errors chain the original provider error for diagnostics. Retryable
408/409 timeouts are not rate-limit waits; caller-level retry handling remains
responsible for body-level non-429 transient failures.
Retaining assistant reasoning¶
OpenRouter and LiteLLM keep separately returned reasoning (including the
reasoning_content alias) and reasoning_details in GenerationResult.metadata.
They do not add it to result.text or to the action produced by
result.to_agent_action(). Structured blocks retain their order, signatures,
nulls and extra provider fields as detached plain data.
For an API continuation, pass that payload in the assistant message:
assistant = ChatMessage(
role="assistant",
content=result.text,
reasoning=result.metadata.get("reasoning"),
reasoning_details=tuple(result.metadata.get("reasoning_details") or ()),
)
ChatMessage.to_dict() sends reasoning_details when nonempty, otherwise
reasoning; it does not send both copies. Messages without either field keep
their ordinary wire representation, and old pickled messages remain readable.
These are provider-native API fields, not a portable prompt format for every
backend. The caller must choose a backend/provider that accepts them and budget
space for the retained reasoning as well as the continuation. Retaining a payload
does not itself enable another reasoning phase or guarantee the provider uses it.
LiteLLM¶
Routes requests through the litellm SDK,
giving unified access to 100+ providers through one interface. Models
use litellm's provider/model format:
from llenvs.inference.backends import LiteLLMBackend
# Direct provider access — API key read from the provider's native
# env var (GEMINI_API_KEY, ANTHROPIC_API_KEY, AZURE_API_KEY, ...)
backend = LiteLLMBackend(model="gemini/gemini-2.5-flash")
# Through a LiteLLM proxy/gateway with a virtual key
backend = LiteLLMBackend(
model="litellm_proxy/Qwen/Qwen3.6-35B-A3B",
api_base="https://your-gateway.example.com",
api_key="sk-your-virtual-key",
)
For the litellm_proxy/ prefix, litellm also reads
LITELLM_PROXY_API_KEY and LITELLM_PROXY_API_BASE natively, so both
constructor params can be omitted.
YAML config:
model:
backend: litellm
model: litellm_proxy/Qwen/Qwen3.6-35B-A3B
max_concurrency: 32
params:
api_base: https://your-gateway.example.com
rate_limit_wait: 30.0
Parameter Dropping¶
drop_params=True (the default) makes litellm silently drop request
parameters the target provider does not support (e.g. presence_penalty
on Anthropic) instead of erroring. Set drop_params=False on the
constructor to fail loudly instead. Logprobs keep a hard guarantee
either way: if logprobs=True was requested and the provider returns
none, the backend raises LogprobsNotReturnedError.
Reasoning / Thinking Budget¶
SamplingParams.thinking_budget maps to litellm's cross-provider
thinking={"type": "enabled", "budget_tokens": N} parameter (Anthropic
thinking.budget_tokens, Gemini thinking budget, Bedrock/Vertex
equivalents); providers without a budget concept drop it via
drop_params. disable_thinking=True sends reasoning_effort="none"
and wins over any other reasoning knob. For effort-style control, pass
extra={"reasoning_effort": "high"} on SamplingParams — params.extra
keys flow as top-level kwargs into litellm.completion() and override
derived values. The logits-level thinking knobs
(thinking_budget_per_block, thinking_budget_suffix,
thinking_budget_soft_ratio) are inert on this backend — they only take
effect for vllm and huggingface.
Reasoning diagnostics land in GenerationResult.metadata:
completion_tokens_details, reasoning_tokens, and presence/length
fields for reasoning_content. When litellm computes a request cost it
is surfaced as metadata["response_cost"].
Rate-Limit Retries¶
Same knobs as OpenRouter: rate_limit_wait (seconds to wait before
retrying a rate-limited call; 0 disables backend-level retries) and
rate_limit_max_retries. The retry path also covers providers that
surface their own 429 as an HTTP-200 body with choices=null and a
rate-limit message. litellm's in-SDK retries are available via
num_retries but default to off — the evaluation runner already
retries transient failures per slot.
Codex CLI¶
from llenvs.inference.backends import CodexCLIBackend
backend = CodexCLIBackend(
model="codex-mini-latest",
max_concurrency=4, # Recommended: keep this small; each request spawns a CLI process
timeout=600.0,
profile="default",
)
Notes¶
- Runs
codex execin a fresh temporary directory for every request. - Locks Codex to
--sandbox read-onlyand disables session persistence with--ephemeral. - Supports chat generation only. Native tool calling, vision, logprobs, prefix continuation, and full scoring are not available.
SamplingParams.max_tokensis accepted but ignored becausecodex execdoes not expose a direct equivalent.- Use the backend constructor's
config_overridesargument to pass Codex-specific-c key=valueoverrides when your local Codex CLI/config supports them.
Capabilities¶
| Feature | Supported | Notes |
|---|---|---|
| logprobs | ✗ | Not exposed by the CLI |
| prefix_continuation | ✗ | Stateless transcript replay only |
| batching | ✓ | Parallel subprocess calls |
| streaming | ✗ | Returns one final assistant message |
| chat | ✓ | Transcript rendered into each request |
| function_calling | ✗ | Native tool calls unsupported |
| vision | ✗ | Text-only |
Claude Code CLI¶
from llenvs.inference.backends import ClaudeCodeBackend
backend = ClaudeCodeBackend(
model="sonnet", # alias, or a full name like "claude-sonnet-4-6"
max_concurrency=4, # Recommended: keep this small; each request spawns a CLI process
timeout=600.0,
effort="low", # "low" | "medium" | "high" | "xhigh" | "max"
max_budget_usd=0.10, # per-call USD safety cap
)
Notes¶
- Runs
claude --print --output-format jsonin a fresh temporary directory for every request. - Always passes
--no-session-persistenceand--strict-mcp-config. By default--tools ""is sent so the spawned CLI cannot run Bash/Edit/Read tools — the backend behaves as a pure text-completion wrapper. --bareis off by default so the CLI can use the same OAuth/keychain login as the interactive TUI (e.g. a Claude Pro/Max subscription). Passbare=Trueto forceANTHROPIC_API_KEY/apiKeyHelperauth and skip auto-discovery — recommended for reproducible CI.- Supports chat generation only. Native tool calling, vision, logprobs, prefix continuation, and full scoring are not available.
SamplingParams.max_tokensis accepted but ignored because the Claude Code CLI does not expose a direct equivalent. Useeffortto control reasoning level andmax_budget_usdas a per-call USD cap.- Errors from the underlying Anthropic API are classified into
PromptTooLongError(context-limit messages),RuntimeError(auth failures — not retryable), orQuotaExhaustedError(everything else, includingis_error: truepayloads, rate limits, server overload, timeouts, and unparseable output). The CLI itself retries 10× internally before surfacing errors (configurable viaCLAUDE_CODE_MAX_RETRIES). - Use the backend constructor's
extra_argsargument (a list of strings) to pass any Claude Code flag this backend does not model directly.
Capabilities¶
| Feature | Supported | Notes |
|---|---|---|
| logprobs | ✗ | Not exposed by the CLI |
| prefix_continuation | ✗ | Stateless transcript replay only |
| batching | ✓ | Parallel subprocess calls |
| streaming | ✗ | Returns one final assistant message |
| chat | ✓ | Transcript rendered into each request |
| function_calling | ✗ | Native tool calls unsupported |
| vision | ✗ | Text-only |
Sampling Parameters¶
from llenvs.inference import SamplingParams
params = SamplingParams(
max_tokens=2048,
temperature=0.0,
top_p=1.0,
top_k=0,
stop_sequences=("</answer>",),
presence_penalty=0.0,
frequency_penalty=0.0,
n=1,
logprobs=False,
)
Backend-Specific Parameters¶
Use the extra parameter to pass backend-specific options that aren't part of the common interface. These are forwarded directly to the underlying library.
# HuggingFace-specific (transformers)
params = SamplingParams(
temperature=0.7,
extra={
"repetition_penalty": 1.2,
"num_beams": 4,
"do_sample": True,
"length_penalty": 1.0,
},
)
# vLLM-specific
params = SamplingParams(
temperature=0.7,
extra={
"best_of": 3,
"use_beam_search": True,
"length_penalty": 1.0,
},
)
# OpenAI-specific
params = SamplingParams(
temperature=0.7,
extra={
"seed": 42,
"response_format": {"type": "json_object"},
"logit_bias": {123: -100},
},
)
# Anthropic-specific
params = SamplingParams(
temperature=0.7,
extra={
"metadata": {"user_id": "user-123"},
},
)
Note: Parameters in
extratake precedence over the common parameters if there's a conflict. Unknown parameters may cause errors from the underlying API.
Chat Template Parameters¶
Local backends (vLLM and HuggingFace) support passing extra keyword arguments to tokenizer.apply_chat_template() via chat_template_kwargs. This is useful for models like Qwen3 that use special template parameters (e.g., enable_thinking=True).
from llenvs.inference.backends import VLLMBackend
backend = VLLMBackend(
model_path="Qwen/Qwen3-8B",
chat_template_kwargs={"enable_thinking": True},
)
The kwargs are spread into every apply_chat_template() call — both generate_chat() and generate_chat_batch().
Thinking Budget¶
This section describes the in-process logits-processor implementation
used by vllm and huggingface — those backends own token generation
and can intervene per-step. For OpenRouter the same thinking_budget
field is accepted but flows through the provider's reasoning API
(reasoning.max_tokens); see the OpenRouter section above for the
behavior and limitations there. Pure HTTP backends without a reasoning
API (OpenAI, Anthropic direct, Codex CLI, Claude Code CLI, vllm_singularity) ignore
these fields entirely.
For models that produce <think>...</think> reasoning blocks (e.g., Qwen3 with enable_thinking=True), you can cap the number of tokens generated inside each thinking block using thinking_budget on InferenceConfig:
from llenvs.core.config import InferenceConfig
config = InferenceConfig(
max_tokens=4096,
thinking_budget=512,
)
When the budget is reached, the processor forces the model to close the thinking block and continue with its answer. By default, the budget is shared across all <think> blocks — if the model opens multiple thinking blocks, their token counts accumulate toward the same limit.
Per-Block Budget¶
To give each <think> block its own independent budget (counter resets on each new block), set thinking_budget_per_block:
Early Stopping Text¶
When the budget is hit, the processor forces a natural-language suffix before </think> to help the model transition cleanly from reasoning to answering. This is controlled via thinking_budget_suffix. By default it is None (bare </think>). Pass DEFAULT_EARLY_STOPPING_SUFFIX to enable the built-in natural-language transition text.
from llenvs.core.config import InferenceConfig
from llenvs.inference import DEFAULT_EARLY_STOPPING_SUFFIX
# Default: bare </think> (no suffix)
config = InferenceConfig(
max_tokens=4096,
thinking_budget=512,
)
# Enable natural-language transition suffix
config = InferenceConfig(
max_tokens=4096,
thinking_budget=512,
thinking_budget_suffix=DEFAULT_EARLY_STOPPING_SUFFIX,
)
# Custom suffix
config = InferenceConfig(
max_tokens=4096,
thinking_budget=512,
thinking_budget_suffix="\n\nLet me answer directly.\n</think>\n\n",
)
# Default: bare </think> (no suffix)
inference:
max_tokens: 4096
thinking_budget: 512
# Custom suffix
inference:
max_tokens: 4096
thinking_budget: 512
thinking_budget_suffix: "\n\nLet me answer directly.\n</think>\n\n"
The suffix should end with </think> so the model naturally transitions out of the thinking block.
Soft Budget Transition¶
The hard cutoff can cause reasoning to leak into the response when the model is mid-thought. To mitigate this, set thinking_budget_soft_ratio to begin boosting the </think> logit before the hard cutoff:
With soft_ratio=0.9 and budget=512, the </think> logit boost begins at token 460 (90% of 512) and ramps linearly up to a max boost of ~5.0 at token 511. At token 512 the hard cutoff forces </think> as before. This gives the model a chance to close its thinking block naturally before being forced.
The feature is opt-in — omitting thinking_budget_soft_ratio preserves the hard-cutoff-only behavior. The ratio must be between 0 and 1 (exclusive).
Requirements:
- The tokenizer must have <think> and </think> as single dedicated tokens
- Works with both vLLM (logits_processors) and HuggingFace (logits_processor) backends
- The vLLM processor is stateful (O(1) per token); the HuggingFace processor derives state from the full token history for batch safety
- vLLM V1 support: On vLLM's V1 engine (default since 0.8), the thinking budget processor is registered as an AdapterLogitsProcessor at engine init, and per-request budgets are passed via SamplingParams.extra_args. On V0, per-request logits_processors callables are used instead. Both paths are transparent — just set the thinking_budget* fields on InferenceConfig
Second Elicitation¶
When a generation hits the max_tokens limit, the output is abruptly truncated — often mid-thought for reasoning models, never producing an answer. Second elicitation performs a follow-up generation with the truncated output in context, using a smaller token budget, to coax a final answer.
Disabled by default. When enabled, it only triggers for generations with finish_reason == MAX_TOKENS.
Configuration¶
from llenvs.core.config import InferenceConfig
from llenvs.inference import DEFAULT_EARLY_STOPPING_SUFFIX
config = InferenceConfig(
max_tokens=2048,
second_elicitation_suffix=DEFAULT_EARLY_STOPPING_SUFFIX, # enables the feature
second_elicitation_max_tokens=256, # budget for the follow-up call
)
inference:
max_tokens: 2048
second_elicitation_suffix: "\n\nPlease answer now.\n" # any non-null value enables the feature
second_elicitation_max_tokens: 256
When a generation is truncated (MAX_TOKENS), the runner:
- Appends the truncated output + suffix as an assistant message
- Appends
"Please provide the final answer now. Follow the formatting instructions specified above exactly."as a user message - Calls the backend with
second_elicitation_max_tokensas the token budget - Merges the two results: concatenated text, summed token counts, second call's finish reason
The merged result's metadata includes "second_elicitation": True.
Custom Suffix¶
The suffix you set on second_elicitation_suffix is appended to the truncated output in the follow-up call. DEFAULT_EARLY_STOPPING_SUFFIX provides a natural-language transition that helps the model shift from reasoning to answering. To customize:
max_model_len Sizing¶
For vLLM, your max_model_len must accommodate the full sequence: prompt + max_tokens + suffix + second_elicitation_max_tokens. The runner does not auto-adjust max_model_len — set it via ModelConfig.params:
model:
backend: vllm
model: Qwen/Qwen3-8B
params:
max_model_len: 8192 # Enough for prompt + first gen + elicitation
Interaction with Thinking Budget¶
Second elicitation and thinking budget are complementary for the first generation, but the follow-up call is forced into no-thinking mode. The runner clears local thinking-budget fields, sets SamplingParams.disable_thinking=True, and backends that support per-call thinking control honor that request:
- vLLM and HuggingFace call the chat template with
enable_thinking=Falseand skip thinking-budget logits processors. - The Singularity-hosted vLLM server backend injects
chat_template_kwargs={"enable_thinking": false}into the request body (the server renders the chat template, so the toggle must travel with the request). - OpenRouter sends
extra_body.reasoning = {"effort": "none"}while preserving unrelated request body fields such as provider routing.
This keeps the rescue call focused on producing the final answer in the required format instead of spending its smaller token budget on another reasoning trace.
Batch Generation¶
All backends support generate_chat_batch() for processing multiple independent conversations efficiently. This is the foundation for parallel evaluation.
from llenvs.inference import ChatMessage, SamplingParams
messages_batch = [
[ChatMessage(role="user", content="What is 2+2?")],
[ChatMessage(role="user", content="What is 3*3?")],
[ChatMessage(role="user", content="What is 10/2?")],
]
results = backend.generate_chat_batch(messages_batch, SamplingParams())
for result in results:
print(result.text)
How each backend implements batching:
- vLLM / HuggingFace: Converts all conversations to prompt strings and calls the model in a single batched GPU call. Maximum throughput.
- API backends (OpenAI, Anthropic, OpenRouter): Fires concurrent async HTTP requests with a semaphore limiting parallelism to
max_concurrency. No API provider supports sending multiple prompts in a single call, so concurrency is the only way to parallelize.
Tool-calling has an equivalent generate_with_tools_batch().
Capabilities¶
Check what a backend supports:
caps = backend.capabilities
print(f"Logprobs: {caps.supports_logprobs}")
print(f"Batching: {caps.supports_batching}")
print(f"Prefix continuation: {caps.supports_prefix_continuation}")
print(f"Function calling: {caps.supports_function_calling}")
Prompt Engineering¶
For pre-built system prompts, question templates, and model profiles, see the Prompts guide. The pipeline system below is the low-level API that those higher-level tools build on.
Building Pipelines¶
from llenvs.inference.prompting import (
AnswerFormatInjector,
ChainOfThoughtWrapper,
FewShotInjector,
SystemPromptInjector,
)
# Compose with >> operator
pipeline = (
SystemPromptInjector("You are an expert mathematician.")
>> FewShotInjector(
[
("What is 2+3?", "2 + 3 = 5\n<answer>5</answer>"),
]
)
>> ChainOfThoughtWrapper("think_step_by_step")
>> AnswerFormatInjector("xml_answer", tag_name="answer")
)
# Apply to messages
messages = [ChatMessage(role="user", content="What is 7*8?")]
transformed = pipeline.transform(messages)
Standard Pipeline Helper¶
from llenvs.inference.prompting import build_standard_pipeline
pipeline = build_standard_pipeline(
system_prompt="You are a helpful assistant.",
examples=[("Q1", "A1"), ("Q2", "A2")],
use_cot=True,
answer_format="xml_answer",
tag_name="answer",
)
Chain of Thought Styles¶
from llenvs.inference.prompting import ChainOfThoughtWrapper
ChainOfThoughtWrapper("think_step_by_step")
# → "Think through this step by step..."
ChainOfThoughtWrapper("show_work")
# → "Show your work and reasoning..."
ChainOfThoughtWrapper("explain")
# → "Explain your thought process..."
Answer Format Options¶
from llenvs.inference.prompting import AnswerFormatInjector
AnswerFormatInjector("xml_answer", tag_name="answer")
# → "Put your final answer in <answer>...</answer> tags."
AnswerFormatInjector("json")
# → "Provide your answer as a JSON object..."
AnswerFormatInjector("boxed")
# → "Put your final answer in \\boxed{}."
AnswerFormatInjector("gsm8k")
# → "End your response with '#### ' followed by..."
Answer Extraction¶
Built-in Extractors¶
from llenvs.core.extraction import (
CompositeExtractor,
GSM8KExtractor,
MultipleChoiceExtractor,
RegexExtractor,
TagBasedExtractor,
)
# XML tags
extractor = TagBasedExtractor(tag_name="answer")
answer, _ = extractor.extract("Result: <answer>42</answer>")
# answer = "42"
# GSM8K format
extractor = GSM8KExtractor()
answer, _ = extractor.extract("So the answer is #### 42")
# answer = "42"
# Multiple choice
extractor = MultipleChoiceExtractor(choices="ABCD")
answer, _ = extractor.extract("The answer is (B)")
# answer = "B"
# Try multiple extractors
extractor = CompositeExtractor(
extractors=[
TagBasedExtractor(tag_name="answer"),
GSM8KExtractor(),
RegexExtractor(pattern=r"(\d+)$"),
]
)