Choose an OpenAI-compatible model upstream¶
FastFence can send protected business completions to native Ollama or to an OpenAI-compatible Chat Completions server. The model ID in your request must be allowed by your policy and available at the selected upstream.
Laya remains independent: security assessment still uses the local Ollama origin
in FASTFENCE_OLLAMA_URL. Choosing a different business provider does not disable
input or output inspection.
| Setting | Purpose |
|---|---|
FASTFENCE_MODEL_PROVIDER=ollama |
Default native Ollama business transport. |
FASTFENCE_MODEL_PROVIDER=openai |
Use the compatible /chat/completions transport. |
FASTFENCE_OPENAI_BASE_URL |
Upstream API base including /v1, such as http://127.0.0.1:11434/v1. |
FASTFENCE_OPENAI_API_KEY |
Optional upstream bearer credential, supplied only to the FastFence server. |
FASTFENCE_OLLAMA_URL |
Ollama origin for Laya and native Ollama, default http://127.0.0.1:11434. |
The upstream key is separate from the agent and management credentials used to call FastFence. Keep it in the server's secret environment. Client requests cannot choose a different upstream URL or supply an upstream credential. Remote bases require HTTPS; HTTP is accepted only for loopback addresses. Inline URL credentials, query strings, fragments, redirects and environment proxy settings are rejected or disabled.
Download the complete examples into your installation's examples/ directory and use the activated FastFence virtual environment. Run commands from the installation directory.
Run against Ollama's compatible API¶
After installing the package, run from your installation directory with ollama serve running in another terminal:
fastfence init --anonymization
fastfence setup-laya
ollama pull qwen3:4b
ollama pull qwen3:0.6b
FASTFENCE_MODEL_PROVIDER=openai \
FASTFENCE_OPENAI_BASE_URL=http://127.0.0.1:11434/v1 \
fastfence serve --port 8002
Ollama exposes its compatible Chat Completions route beneath /v1; its native
route remains available independently. See the
Ollama compatibility documentation.
In a second terminal, run this complete protected smoke request. It reads a local agent credential without printing it, requests a 256-token completion, and checks the gateway verdict as well as upstream execution:
python - <<'PY'
import json
import os
from pathlib import Path
import httpx
token = os.environ.get("FASTFENCE_AGENT_TOKEN")
if not token:
path = Path("state/credentials.json")
if not path.exists():
path = Path("state/demo-tokens.json")
credentials = json.loads(path.read_text())
token = credentials.get("local-agent") or credentials["analyst-blue"]
with httpx.Client(timeout=120, trust_env=False) as client:
response = client.post(
"http://127.0.0.1:8002/api/models/complete",
headers={"Authorization": "Bearer " + token},
json={
"model": "qwen3:0.6b",
"prompt": "Say hello in one sentence.",
"max_output_tokens": 256,
},
)
response.raise_for_status()
verdict = response.json()
print(json.dumps({key: verdict[key] for key in (
"decision", "reason", "semantic_input_status", "semantic_output_status",
"upstream_executed", "output",
)}, indent=2))
assert verdict["decision"] in {"allowed", "redacted"}, verdict["reason"]
assert verdict["upstream_executed"]
assert verdict["output"]["text"].strip()
PY
Use FASTFENCE_MODEL_PROVIDER=ollama to return to the native transport. The normal
REST, MCP and compatible client interfaces remain the same. Model classifications
can vary; an HTTP 200 alone does not mean that the gateway allowed a request.
Reasoning models can spend a short completion allowance entirely on reasoning,
returning empty text with finish_reason: length; adjust the allowed token budget
or the upstream model configuration if this occurs.
Use vLLM¶
On a machine with vLLM installed and sufficient resources for the selected model, start a compatible server with a stable model alias:
vllm serve Qwen/Qwen2.5-0.5B-Instruct \
--host 127.0.0.1 --port 8001 --served-model-name business-model
This uses vLLM's Chat Completions server; model and hardware prerequisites are described in the vLLM server documentation.
In FastFence, add business-model to the model allowlist as described below, then
start the gateway:
FASTFENCE_MODEL_PROVIDER=openai \
FASTFENCE_OPENAI_BASE_URL=http://127.0.0.1:8001/v1 \
fastfence serve --port 8002
For an authenticated server, configure its credential using vLLM's --api-key
option or VLLM_API_KEY environment variable, and inject the same value as
FASTFENCE_OPENAI_API_KEY in the gateway environment. For a remote deployment,
use your server's HTTPS API base instead of the loopback URL.
Use llama.cpp¶
With llama-server installed and a compatible instruction-model GGUF file already
available, point LLAMA_MODEL_PATH at that file:
export LLAMA_MODEL_PATH=/absolute/path/to/your-instruct-model.gguf
llama-server --model "$LLAMA_MODEL_PATH" --alias business-model \
--host 127.0.0.1 --port 8080
The alias becomes the API model ID. The server supports compatible Chat Completions; see the llama.cpp server reference.
Start FastFence with:
FASTFENCE_MODEL_PROVIDER=openai \
FASTFENCE_OPENAI_BASE_URL=http://127.0.0.1:8080/v1 \
fastfence serve --port 8002
Allow and call the served model¶
Connect a management identity in the console, open Policies → Edit
configuration, and add the model under models in Advanced configuration. For
the default initialized analyst/operator budgets, this entry permits both roles:
business-model:
roles: [analyst, operator]
max_output_tokens: 256
timeout_ms: 30000
cost_microusd: 0
The advanced editor accepts the complete policy as JSON; this YAML fragment shows
the entry to add beneath models, rather than a standalone replacement policy.
Use roles with budgets in your actual installation. Review the complete change
and activate the next policy version. Then call the executable client:
python examples/protected_request.py \
--url http://127.0.0.1:8002 --model business-model --prompt 'Hello'
The same allowlisted model is available through the compatible SDK client and MCP client.
Compatibility contract and verification¶
The adapter sends one nonstreaming text chat completion with model, messages,
max_tokens, temperature: 0 and optional stop strings. Prompt-only calls become
one user message. The upstream must return one assistant text choice, an explicit
stop or length finish reason and nonnegative integer prompt_tokens,
completion_tokens and total_tokens with a consistent sum. Tool calls,
function calls, refusals, nontext content, missing or excessive usage, encoded or
oversized response bodies and invalid completion states fail closed. There are
no automatic provider retries.
Response JSON is limited to 262,144 bytes and text to 65,536 UTF-8 bytes; the policy's smaller output bound still applies. The entire exchange uses the model policy deadline. This interface does not implement provider tool execution, streaming, multimodal messages or the Responses API.
The actual Ollama /v1 path was checked with qwen3:0.6b: a 256-token request
returned nonempty text, explicit stop and 156 total provider token units. The
test suite also performs real loopback HTTP exchanges with synthetic replies and
adversarial transport fixtures. vLLM and llama.cpp commands were checked against
their official documentation; those engines were not run during this validation.
Developer transport regression commands are in Contributing.