Quickstart
Create an API key in the Privateer app under Settings → API keys, then set it as your OpenAI base URL and key. The key looks like sk-priv-… and is shown only once.
from openai import OpenAI
client = OpenAI(
base_url="https://api.privateer.pro/v1",
api_key="sk-priv-…",
)
resp = client.chat.completions.create(
model="", # "" = your account's default model; or an id from /v1/models
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.privateer.pro/v1",
apiKey: process.env.PRIVATEER_API_KEY,
});
const resp = await client.chat.completions.create({
model: "",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(resp.choices[0].message.content);
curl https://api.privateer.pro/v1/chat/completions \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"model":"","messages":[{"role":"user","content":"Hello!"}]}'
Authentication
Every request needs a Privateer API key in the Authorization header:
Authorization: Bearer sk-priv-…
Create, name, and revoke keys in the Privateer app under Settings → API keys. The full key is shown once at creation — we store only a hash, so if you lose it, revoke it and mint a new one. A key inherits its owner's plan, rate limits, and credit balance; revoking it stops every request using it immediately.
Pricing. API usage is pay-as-you-go, drawn from your account credit — across chat, image, video, audio and 3D. Per-model rates for the whole catalogue are at privateer.pro/models; how each kind is metered is at what media costs, and how billing works at privateer.pro/api-pricing.
Keys are managed in the app, not over this API. There is deliberately no /v1 endpoint that mints or lists keys — a key cannot create another key, so a leaked one cannot quietly grow itself a replacement. Minting and revoking happen in a signed-in session, in the app.
Keep keys server-side. Anyone holding a key can spend your account balance. Never ship one in a mobile app, browser bundle, or public repo.
Privacy & encryption
In the Privateer app, your content is encrypted on your device or in our cloud, with inference on hardware-attested enclaves. This API is different by nature: you send prompts to our server, which forwards them to the model provider to run inference. That traffic is processed in the clear and is not encrypted.
What we actually guarantee is retention, not secrecy in transit. Inference is a stateless pass-through: we persist only billing metadata (timestamps, model id, token counts, cost) — never your prompts or the responses. By default requests are pinned to Zero-Data-Retention providers, and confidential models (NEAR AI, Tinfoil) run inside Trusted Execution Environments. Set "requireZdr": false in the request body to opt out of ZDR pinning.
Over HTTP, enclave routing is asserted by us — not verified by you. The app earns its Verified badge by checking a TEE attestation on the device, on every response. An API client does no such check, so treat confidential routing here as our policy rather than as proof. If cryptographic verification is what you need, the app and the CLI are the surfaces that can give it to you.
Endpoints
All paths are relative to https://api.privateer.pro/v1.
POST /v1/chat/completions
Standard OpenAI Chat Completions. Tool calls, multi-part content, and finish_reason pass through unchanged. Common parameters:
| Field | Type | Notes |
|---|---|---|
messages | array | Required. The conversation, OpenAI format. |
model | string | An id from /v1/models. Empty or omitted uses your account's default model. |
stream | boolean | When true, responses stream as SSE (see below). |
max_tokens, temperature, … | — | Passed through to the provider unchanged. |
requireZdr | boolean | Privateer extension. Defaults to your account setting (ZDR on). Set false to allow non-ZDR routing. |
A successful response is the provider's standard Chat Completion object, including a usage block that determines billing.
Streaming
Set "stream": true to receive Server-Sent Events. Each event is a chat.completion.chunk; the stream ends with data: [DONE]. The final chunk carries usage.
curl -N https://api.privateer.pro/v1/chat/completions \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"model":"","stream":true,"messages":[{"role":"user","content":"Count to 3"}]}'
With the OpenAI SDKs, pass stream=True (Python) / stream: true (Node) and iterate the returned stream as usual.
Vision (image input)
Image input works on /v1/chat/completions with no extra endpoint — pass OpenAI multi-part content and a vision-capable model:
curl https://api.privateer.pro/v1/chat/completions \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{
"model": "near/Qwen/Qwen3.5-122B-A10B",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,…"}}
]
}]
}'
Pass a model whose id supports image input (see Models). Data-URI and remote image_urls are both accepted.
Models
GET /v1/models returns the models available to your account in OpenAI's list shape. Use any returned id as the model field; pass an empty string to use your account default.
curl https://api.privateer.pro/v1/models \
-H "Authorization: Bearer sk-priv-…"
{
"object": "list",
"data": [
{ "id": "near/deepseek-ai/DeepSeek-V4-Flash", "object": "model", "owned_by": "privateer" }
]
}
/v1/models lists chat models only. Image, video, audio and 3D models are not chat completions and are not in that response — each has its own catalogue: GET /v1/models3d and GET /v1/sprites/models on the API, and privateer.pro/models for the full priced list of every one of them, ids included.
What media costs
Chat is billed per token. Everything else is billed per unit of output, and the unit is the provider's own — we don't convert, because converting would mean inventing the resolution, duration or token count your request actually decides.
| Endpoint | Metered by |
|---|---|
/v1/chat/completions | Input and output tokens, priced separately. |
/v1/images/generations | Per generated image, per megapixel, or per output token — whichever the model's provider meters. Billed per image produced, so a partial n is never free and never overcharged. |
/v1/videos | Per second of output, which usually varies by resolution; a few models meter per video token instead. Reserved at submit, settled against the provider's actual cost on completion. |
/v1/audio/transcriptions | Per minute of audio submitted. |
/v1/audio/speech | Per 1,000 characters of input, counted after the 8,000-character cap is applied — you are billed for what we send, not what you sent. |
/v1/audio/sfx | Flat per generation, or per second of the duration you asked for, depending on the model. |
/v1/audio/music | Flat per track, or per minute of the duration you asked for, depending on the model — so a five-minute track and a thirty-second clip are not the same price. Leaving the prompt rewrite on adds one small chat call. |
/v1/models3d | Per mesh, and the price moves with the axes you pick — so GET /v1/models3d returns the price of every option, and a submit quotes your exact request before you commit. |
/v1/sprites/generate | One video generation per billed facing, at the video rate. /v1/sprites calls no model and is not billed at all. |
Live per-model rates for all of them — image, video, music, sound effects, voices and 3D, each with its id and its unit — are on privateer.pro/models, quoted at the same developer-API rate this key bills at.
Image generation
OpenAI Images shape. Returns base64 by default (b64_json), or set response_format:"url" for a short-lived signed URL. n is capped at 4. Leave model empty for your account's default image model.
curl https://api.privateer.pro/v1/images/generations \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"prompt":"an otter surfing, watercolor","n":1}'
{
"created": 1783437233,
"data": [ { "b64_json": "/9j/4AAQ…" } ]
}
| Field | Notes |
|---|---|
prompt | Required. |
model | Image model id, or empty for the account default. |
n | 1–4 images. |
size, aspect_ratio | Optional, passed to the model. |
response_format | b64_json (default) or url. |
Video generation
Video generation is asynchronous: submit a job, then poll it. This is a Privateer extension (no OpenAI equivalent). The finished video is returned as a short-lived signed URL.
curl https://api.privateer.pro/v1/videos \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"prompt":"a butterfly landing on a flower","seconds":4}'
# → 202
{ "id": "NzX8ue0…", "object": "video.job", "status": "queued" }
curl https://api.privateer.pro/v1/videos/NzX8ue0… \
-H "Authorization: Bearer sk-priv-…"
# processing…
{ "id": "NzX8ue0…", "status": "processing" }
# done
{ "id": "NzX8ue0…", "status": "completed",
"url": "https://…signed…", "expires_at": 1783440833 }
This endpoint is text-to-video: the prompt is the only conditioning it takes. Image conditioning — pinned first/last frames, reference images, continuing or altering an existing clip — is available in the Privateer apps, not on /v1. For image conditioning that is on the API, see 3D model generation and sprite sheets below.
Optional submit fields: model, seconds, size, generate_audio, requireZdr. The signed url is short-lived — each poll of a completed job returns a fresh one. Video requires a paid plan (or held pay-as-you-go credit).
Poll every few seconds. Billing is reserved at submit and settled against the provider's actual cost on completion; a failed job releases the hold.
3D model generation
Turn reference images into a 3D mesh you can import into a game engine or DCC tool. Like video, this is asynchronous and a Privateer extension: submit a job, then poll it. The finished mesh is returned as a short-lived signed URL.
It is image-to-mesh — there is no text-to-mesh endpoint. Send between one and four views of the same object as base64 (or a data: URI); extra views stop the model inventing the sides it cannot see. Remote image URLs are not accepted.
Several models are available, from different labs, and they do not take the same options — one is configured by output resolution, another by a texture tier, another by a quality preset. So start by listing them: GET /v1/models3d returns every model with its price, its output formats, how many views it uses, and the exact options it accepts. Pass the ones you choose as axes.
curl https://api.privateer.pro/v1/models3d \
-H "Authorization: Bearer sk-priv-…"
{ "object": "list", "data": [
{ "id": "fal-ai/trellis-2", "object": "model3d", "name": "Trellis 2",
"formats": ["glb"], "max_views": 1, "price_usd": { "min": 0.35, "max": 0.49 },
"axes": [
{ "name": "resolution", "kind": "enum", "priced": true,
"default": "1024", "values": ["512", "1024", "1536"] },
{ "name": "faceCount", "kind": "int", "priced": false,
"min": 5000, "max": 2000000 } ],
"conflicts": [] }, … ] }
curl https://api.privateer.pro/v1/models3d \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"image":"iVBORw0KGgo…","model":"fal-ai/trellis-2",
"axes":{"resolution":"1536","faceCount":250000}}'
# → 202
{ "id": "01a0051b…", "object": "model3d.job", "status": "queued",
"format": "glb", "estimated_cost_usd": 0.49 }
curl https://api.privateer.pro/v1/models3d/01a0051b… \
-H "Authorization: Bearer sk-priv-…"
# processing… (typically 60–90 seconds)
{ "id": "01a0051b…", "status": "processing" }
# done
{ "id": "01a0051b…", "status": "completed", "format": "glb",
"url": "https://…signed…", "expires_at": 1783440833 }
| Field | Type | Notes |
|---|---|---|
image / images | string / array | Required. 1–4 base64 images or data: URIs of the same object, read as front, back, left, right. 12 MB total. How many are actually used is per-model — see max_views. |
model | string | An id from GET /v1/models3d. Omit for your account's default. |
axes | object | The chosen model's own options, as {"name": value}. Names and legal values come from that model's axes in the list response; an option the model doesn't have is rejected rather than ignored. Omit one to take its default — for a numeric budget that means letting the provider choose, which on some models is what avoids a surcharge. |
format | string | One of the model's formats; the first is the default. Prefer glb — it is the only container that carries the materials. |
generate_type | string | Hunyuan models only, and equivalent to axes.generateType. Normal (default, textured), LowPoly (retopologised, best for anything that deforms), Geometry (untextured, for blockouts). |
polygon_type | string | Hunyuan models only. triangle (default) or quadrilateral. Quads deform better under animation. |
face_count | integer | Target face budget, honoured exactly; the legal range is per-model. Omit for the provider's default, which produces a much larger file. |
pbr | boolean | Generate PBR materials (base colour, normal, metallic/roughness). Rejected on any model producing untextured geometry. |
The named fields above are the older, Hunyuan-shaped spelling and still work; axes is the general one and wins where both are sent. It is the only way to reach the options on the other models.
Cost depends on which model you pick and which options you set — the spread across the catalog is more than tenfold, and an option that is a surcharge on one model is free on another, which is why the price lives in the list response rather than here. The submit response reports estimated_cost_usd for your exact request before you commit to polling; that is the figure that gets charged. 3D generation requires a paid plan (or held pay-as-you-go credit).
Poll every few seconds. Billing is reserved at submit and settled on completion; a failed job releases the hold. Each poll of a completed job returns a fresh signed URL until the mesh is swept from storage.
Sprite sheets (Godot)
Turn frames into a packed sprite sheet with a Godot 4 SpriteFrames resource beside it, or start a facing further back than that — a picture and a sentence — and let Privateer render the motion first. Either way the response is one archive: the sheet, the individual frames, a .tres you can drop straight onto an AnimatedSprite2D, and a README saying where to put them.
Two endpoints, because they need very different things from the deployment:
POST /v1/spritespacks frames you already hold. It calls no model, so it is synchronous and not billed — there is no daily cap and no balance check on it.POST /v1/sprites/generaterenders the frames first, from one image and a description of the motion. It is asynchronous and billed, one video generation per facing that isn't mirrored.
Start at GET /v1/sprites/models. It is the only way to tell “this deployment can't” from “my request was wrong”: generation needs a video decoder on the server, and where there isn't one, generate.available is false and generation refuses with 503 SPRITE_GENERATION_UNAVAILABLE before anything is billed. Packing is unaffected and works everywhere. Like the 3D catalog, this endpoint is behind no plan or balance gate.
curl https://api.privateer.pro/v1/sprites/models \
-H "Authorization: Bearer sk-priv-…"
{ "object": "list",
"pack": { "available": true, "input": "PNG frames, base64" },
"generate": { "available": true, "decoder": "ffmpeg 7.1" },
"direction_sets": [
{ "id": "one", "directions": ["default"], "billed": 1, "mirrored": 0 },
{ "id": "four", "directions": ["down", "right", "up", "left"], "billed": 3, "mirrored": 1 },
{ "id": "eight", "directions": ["down", "down_right", …], "billed": 5, "mirrored": 3 } ],
"limits": { "frames": { "min": 2, "max": 24 },
"frame_size": { "min": 8, "max": 512 },
"max_facings": 8, "max_frame_bytes": 8388608,
"max_total_input_bytes": 50331648 },
"output": { "godot": "4.x", "resource": "SpriteFrames", "container": "application/zip" } }
Packing frames you already have — POST /v1/sprites
Send base64 PNGs. One facing is a flat frames array; several are facings, each with its own direction. Every generated facing must carry the same number of frames — the sheet is a grid, and a short row would leave dead cells the resource still points at.
curl https://api.privateer.pro/v1/sprites \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"name":"knight","action":"walk","directions":"four","fps":12,
"frame_size":64,
"facings":[{"direction":"down","frames":["iVBORw0…","iVBORw0…"]},
{"direction":"right","frames":["iVBORw0…","iVBORw0…"]},
{"direction":"up","frames":["iVBORw0…","iVBORw0…"]},
{"direction":"left","mirror_of":"right"}]}'
A facing with mirror_of costs nothing and sends no frames — it is the named facing flipped horizontally, which is how left comes free from right. It must name a facing that is in the same request.
| Field | Type | Notes |
|---|---|---|
frames / facings | array | Required. Base64 PNGs (or data: URIs). 2–24 frames per facing, up to 8 facings, 8 MB per frame and 48 MB per request. PNG only, in and out — JPEG cannot carry the alpha that keying exists to produce. |
name, action | string | The sheet's name and the animation's. Animations come back named action_direction — walk_down, walk_left — which is what GDScript plays. |
directions | string | one, four or eight. Sets the facing vocabulary and the sheet's row order. |
frame_size | integer | Cell edge in pixels, 8–512, default 64. Frames are cropped to one union box across every facing, then scaled down nearest-neighbour, so the character stays registered instead of juddering in place. |
fps, loop | integer, boolean | Written into the resource. Default 12 and true. Retiming later rewrites one float in the .tres. |
chroma | object / false | Backdrop to key out: { "key": "#00FF00", "tolerance": 0.22, "softness": 0.12, "despill": 0.8 }. Send false if your frames already carry alpha — keying them again would eat whatever their own backdrop is. |
include_frames | boolean | Default true: the archive also carries each cell as its own PNG. |
res_path | string | Where the files will live in your project, default res://sprites/<name>/. It is what the .tres points at, so getting it right here saves an edit. |
# → 200
{ "id": "sprite_9f2c…", "object": "sprite", "status": "completed",
"url": "https://…signed…", "expires_at": 1783440833, "bytes": 148213,
"sheet": { "width": 512, "height": 256, "columns": 8, "rows": 4,
"frame_width": 64, "frame_height": 64 },
"animations": [ { "name": "walk_down", "direction": "down", "origin": "generated" },
{ "name": "walk_left", "direction": "left", "origin": "mirrored" }, … ],
"fps": 12, "loop": true, "res_path": "res://sprites/knight/",
"spriteframes_tres": "[gd_resource type=\"SpriteFrames\" …",
"key_residue": 0.0031 }
The archive at url holds <name>_sheet.png, the .tres, a README, and (unless you turned them off) frames/. The resource is also returned inline as spriteframes_tres: it is a few kilobytes of text, and a build step very often wants to read or rewrite the res:// path without unzipping anything.
key_residue is the honest-failure signal, surfaced rather than swallowed: it is how much backdrop survived the key. A flat backdrop is something you asked a model for and it can ignore, and when it does you get a rim. Raise chroma.tolerance and pack again — re-packing costs nothing. If the key went the other way and removed the subject with the backdrop, the request fails with 422 SPRITE_KEY_EMPTY rather than handing you a sheet of speckle.
Generating the frames — POST /v1/sprites/generate
One image and one sentence. Privateer turns your picture to face each direction, renders a short looping clip per facing against a flat backdrop, samples the frames out of it, and packs them exactly as above. Submit, then poll GET /v1/sprites/{id}.
It is image-to-video and refuses without an image rather than quietly falling back to text: a text-rendered clip of “a knight walking” is a different knight every facing, and the turned facings are edits of the picture you supplied — there is nothing to turn without one.
curl https://api.privateer.pro/v1/sprites/generate \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"image":"iVBORw0KGgo…","prompt":"walking, steady gait, looping",
"name":"knight","action":"walk","directions":"eight",
"frames":8,"frame_size":64,"fps":12}'
# → 202
{ "id": "sprite_9f2c…", "object": "sprite.job", "status": "queued",
"billed_facings": 5, "mirrored_facings": 3, "animations": 8 }
# poll
curl https://api.privateer.pro/v1/sprites/sprite_9f2c… \
-H "Authorization: Bearer sk-priv-…"
{ "id": "sprite_9f2c…", "status": "processing",
"facings_ready": 2, "facings_total": 5 }
# done — the same body POST /v1/sprites returns, plus object: "sprite.job"
Mirroring is the cost story, and it is said at submit. billed_facings is what you pay for; animations is what you get. An eight-way set costs five video generations, not eight, because left, up_left and down_left are their right-hand counterparts flipped — and the flip happens after keying, so the alpha is byte-identical.
Everything in the packing table above applies here too, plus prompt (required — describe the motion), image (required, base64 or a data: URI), model (a video model id), and image_model (the model that draws the turned facings). Nothing packs until every billed facing has landed: the union bounding box that keeps the character registered is computed across all of them. A facing that fails is dropped and the rest still make a sheet. Billing is reserved per clip at submit and settled on completion; a failed clip releases its hold, and an underfunded fan-out releases everything it had reserved rather than holding your balance against clips that can never be packed.
Generation is behind the non-ZDR media gate, like sound effects and 3D: the video and image models it drives have no Zero-Data-Retention endpoint, so a default account gets 403 ZDR_MEDIA_BLOCKED until non-ZDR media generation is enabled — per request with "allowNonZdrMedia": true, or as an account setting. This is the same rule the app enforces. POST /v1/sprites calls no model and is not gated.
Audio (speech-to-text, text-to-speech, sound effects & music)
Four endpoints, each taking an optional model. There is no /v1/audio/models to list from: the voices, the sound-effect generators, the music models and the transcription models each come from a different catalogue, and all of them — ids, prices and the unit each is metered in — are published on privateer.pro/models. Omit model and your account's default is used.
Transcription (STT) — POST /v1/audio/transcriptions
OpenAI Whisper shape: multipart upload with a file field (or JSON { "audioBase64", "format" }). Returns { "text": … }.
curl https://api.privateer.pro/v1/audio/transcriptions \
-H "Authorization: Bearer sk-priv-…" \
-F file=@speech.mp3 \
-F model=openai/whisper-1 \
-F requireZdr=false
Speech (TTS) — POST /v1/audio/speech
Returns raw audio bytes. The default voice model returns audio/wav; request response_format:"mp3" only with an mp3-capable model. input is capped at 8,000 characters — longer text is truncated rather than rejected, and you are billed for what we send.
curl https://api.privateer.pro/v1/audio/speech \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"input":"Hello from Privateer.","voice":"Zephyr"}' \
--output speech.wav
Audio is ZDR-pinned by default, like chat. The default STT model has no Zero-Data-Retention endpoint, so transcription needs requireZdr:false or a confidential (enclave) STT model.
Sound effects — POST /v1/audio/sfx
One short sound from a text description. Returns raw audio/mpeg bytes, plus an X-Privateer-Duration-Seconds header giving the length actually rendered.
input is required (max 500 characters). duration is in seconds and is clamped to 1–30, defaulting to 5 — a value outside that range is corrected, not rejected.
curl https://api.privateer.pro/v1/audio/sfx \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"input":"a heavy wooden door slamming shut","duration":3}' \
--output door.mp3
This is the one audio endpoint behind the non-ZDR media gate. Sound effects run on a provider with no Zero-Data-Retention endpoint, so if your account requires ZDR you'll get 403 ZDR_MEDIA_BLOCKED until non-ZDR media generation is enabled — either per request with "allowNonZdrMedia": true, or as an account setting. This is the same rule the app enforces, not an API-only restriction. Your prompt is sent unattributed and we ask the provider to store neither it nor the generated file, but we can't promise zero retention the way we can for a ZDR-pinned model — so the choice stays yours to make explicitly.
Music — POST /v1/audio/music
A full track — sung or instrumental — from a text brief. Returns raw audio bytes, with the model that rendered it in X-Privateer-Model.
input is required (max 2,000 characters). duration is in seconds and is clamped to what the chosen model will render; several models treat it as a hint rather than a contract, which is why the response carries both X-Privateer-Duration-Seconds (what came back) and X-Privateer-Requested-Seconds (what you asked for) instead of relabelling one as the other. Some models are fixed-length and report neither.
curl https://api.privateer.pro/v1/audio/music \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"input":"slow dub techno, deep sub bass, tape hiss","duration":60}' \
--output track.mp3
| Field | Type | Notes |
|---|---|---|
input | string | Required. The brief. prompt is accepted as an alias. |
model | string | A music model id from privateer.pro/models. Omitted uses your account's default. |
duration | number | Seconds. Clamped to the model's own range rather than rejected; ignored by fixed-length models. |
instrumental | boolean | Defaults to true. Set false for a sung track. |
lyrics | string | The words to sing. Several models require them when instrumental is false and refuse the request up front — before anything is billed — when they are missing or too short. |
refine_prompt | boolean | Defaults to true. See below. |
Your brief is rewritten before it is sent, unless you say otherwise. Music models have no negation (“no vocals” asks for vocals), every provider refuses prompts naming a real artist or track — after the wait, and at full price — and two of these models read tag lists rather than sentences. So a small chat model rewrites the brief into the shape the chosen model reads, as one extra billed chat call. Send "refine_prompt": false if you wrote the brief for a specific endpoint and mean it literally.
This call is synchronous and can take minutes. Unlike video and 3D there is no job to poll — the track comes back in the response, and a long one may take several minutes to render. Raise your client's timeout accordingly: an SDK default of 60 seconds will abandon a track you have already paid for.
Music is not a Zero-Data-Retention model, and it is the one media endpoint here with no gate in front of it. No music generator on any host we reach offers a ZDR or confidential endpoint — so unlike sound effects, there is no other model to switch to, and gating it would mean no music at all rather than a choice. It is exempt on purpose: the call goes through even on an account that requires ZDR, because there is nowhere zero-retention to send it — you will never see 403 ZDR_MEDIA_BLOCKED from this endpoint. What we do instead is send the request unattributed — no account identifier, no conversation, nothing tying the brief to you — and tell you plainly that we cannot promise the provider forgets it. Never treat a music prompt as private. The apps carry this same disclosure at the point of generation.
Search & fetch
The piece an OpenAI-compatible base URL can't supply on its own: live web results to ground an answer with. Same base URL, same sk-priv-… key, same bill — no second vendor and no second API key to hold.
Search — POST /v1/search
query is required. count is clamped to 1–10 (default 5), and freshness takes pd, pw, pm or py — past day, week, month or year — with anything else ignored rather than rejected.
curl https://api.privateer.pro/v1/search \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"query":"webgpu browser support","count":5,"freshness":"pm"}'
Returns { "query", "results": [{ "title", "url", "description", "age", "image" }] }. age and image are present only when the index has them. Add "include_media": true for images and videos arrays alongside the results — they ride along on the same upstream response, so they cost nothing extra and are withheld by default only to keep the payload small.
Fetch — POST /v1/fetch
Reads pages and returns their extracted text. Pass one url, or a urls array of up to 3 — anything beyond that comes back marked failed rather than being silently dropped. http/https only; private and internal addresses are refused.
curl https://api.privateer.pro/v1/fetch \
-H "Authorization: Bearer sk-priv-…" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com/changelog"}'
Returns { "anyFailed", "results": [{ "ok", "url", "title", "text" }] }, with error in place of text on a page that couldn't be read. Text is truncated at 20,000 characters per page. A URL that fails is reported per result — the request itself still succeeds.
What "private" means here. The search provider never sees you: no key of yours, no IP of yours, no account of yours with them, and nothing to join your queries together across requests. We hold the provider key and make the call; page fetches leave from our egress, not your server's. As with the rest of this API, the query itself reaches us in the clear — this is not end-to-end encrypted, and we don't claim it is. What we do guarantee is retention: billing metadata only, never the query, the results, or the page text.
Caps. Search spends your plan's daily web-search allowance and is charged per call at the developer-API rate; a provider failure isn't billed. Fetch costs nothing but spends the daily link-fetch allowance and is bounded by a per-minute egress budget. Both share the standard rate limit and concurrency pool.
Agent Client Protocol
Everything above is the hosted inference API. This section is different: it's about the Privateer agent — the one that runs on your own machine — and how other tools can drive it.
Privateer speaks the Agent Client Protocol (ACP), the same interface Zed and other editors use to talk to coding agents. A host launches privateer acp and speaks newline-delimited JSON-RPC over stdio; the host owns the transport and the interface, Privateer owns the agent.
privateer acp
The point of running it this way is that the moat comes with it. The host renders approvals, it does not grant them:
- The permission gate still decides. Every action is classified locally and gated actions are sent to the host as
session/request_permission. A cancelled dialog, an unparseable answer, or a host that can't be reached all resolve to deny. - The tool ceiling is yours, not the host's. Set
acp.toolsin~/.privateer/config.json. A host cannot widen it. - Destructive actions can never become standing permission. Dangerous shell and guarded files are never offered an "allow always" option, so they re-confirm every time.
- Your models, your account. The agent offers its own model list over the protocol, including the confidential TEE models, and runs on your subscription.
Configure it under an acp block in ~/.privateer/config.json. The default ceiling is read-only:
{
"acp": {
"tools": ["read", "grep", "find", "ls"],
"posture": "approve"
}
}
posture is readonly, approve (default — each risky action asks), or auto. The working directory comes from the host per session, and access outside it is refused rather than prompted.
Add Privateer to Buzz
Buzz is Block's self-hostable workspace where people and agents share the same channels. Because Buzz drives agents over ACP, Privateer drops in as a custom harness and you can message it like a teammate.
In Buzz Desktop, add a custom harness with:
| Field | Value |
|---|---|
| Name | Privateer |
| Command | the absolute path to your node binary, e.g. /usr/local/bin/node |
| Arguments | one argument: the absolute path to bin/privateer-acp.mjs |
Use absolute paths for both. Buzz is a desktop app and inherits a minimal PATH, so a version-managed node (nvm, asdf) usually isn't on it. The path to privateer-acp.mjs is a separate argument, not part of the command.
Then create an agent that uses the harness. Two settings are worth changing:
| Setting | Value | Why |
|---|---|---|
| Parallelism (on the agent, not the harness) | 1 | Privateer holds one account session per process, so parallel copies contend for it. One process also cuts startup from about a minute to a few seconds. Set this on the agent record — Buzz passes it as an explicit flag, and flags beat the BUZZ_ACP_AGENTS env var. |
BUZZ_ACP_LAZY_POOL (env var) | true | Join the channel first and warm the agent pool after, so the agent shows online immediately. It is a boolean — 1 is rejected and the harness exits with status 2. |
Buzz auto-approves every permission request. Its harness answers session/request_permission itself by selecting the "allow once" option — there is no path to a human, and no setting changes that. (Buzz's --permission-mode flag only reaches agents that implement session/set_config_option, which Privateer does not.) Privateer still classifies every action and still refuses anything outside the session's working directory, but in Buzz a posture of approve behaves like auto. Your real controls are acp.tools and acp.posture — a host cannot widen either. Choose the ceiling on the assumption that everything in it will run unattended.
Buzz holds the agent's Nostr identity, so its display name and avatar are set on the Buzz side, not by Privateer.
Connectors (MCP)
There is no MCP endpoint on this API. The hosted API above is inference only — it has no tools and holds no credentials of yours. MCP is a capability of the Privateer agent, the one that runs on your own machine. If you're looking for a way to give your product tools, that's your MCP client's job; this section is about giving the agent tools.
Privateer is an MCP client. Point it at a Model Context Protocol server and that server's tools become first-class agent tools — available to the CLI, the desktop app, the always-on harbor's unattended routine runs, and to any ACP host driving the agent. Two kinds:
| Transport | What it means |
|---|---|
stdio | Local. Privateer spawns the server as a child process on your machine. Nothing leaves the box except what that server itself chooses to send. |
http | Remote. An https endpoint someone else hosts, authenticating with oauth (you authorize in a browser on that machine), a static bearer token, or nothing. Whatever the agent hands it leaves your machine. |
Configuration
Connectors live in two files under the agent's home. The first is the source of truth; the second is generated from it and is the standard shape the MCP adapter reads — edit the source, never the projection.
~/.privateer/agent/mcp-desktop.json # source of truth — every connector, each with `enabled`
~/.privateer/agent/mcp.json # projection: enabled connectors only
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "…" }
},
"linear": {
"url": "https://mcp.linear.app/sse",
"auth": "oauth"
}
},
"settings": { "toolPrefix": "server" }
}
Both files sit in the shared ~/.privateer home, so one machine has one coherent connector config whichever surface edited it — the terminal, the desktop app over loopback IPC, or the phone over the relay.
Adding one
In the terminal, /connect opens a picker over a curated catalog of 22 connectors (GitHub, Slack, Notion, Linear, Jira & Confluence, Sentry, Stripe, Asana, Supabase, Figma, Gmail, Google Drive, PostgreSQL, Playwright, Filesystem, …) plus a Custom connector entry that takes any stdio command line or any https:// URL. It reloads the adapter in place, so new tools are live in the session you're already in. /mcp is the adapter's own status view — what actually connected. The Privateer app and the desktop app have the same editor.
Tools and the permission gate
By default the adapter exposes MCP through a single proxy tool named mcp: one grant covers every enabled server. When a scheduled routine carries a per-connector allow-list, Privateer instead scopes that run to exactly the selected servers and tools, each registered under its own <server>_<tool> name — so an unattended task can hold GitHub's create_issue without holding all of MCP.
Either way, every MCP tool is classified and gated exactly like a built-in. Under ACP that means an MCP tool call surfaces as a session/request_permission to the host, and your acp.tools ceiling still applies — a tool is not trusted because you configured the server it came from.
Credentials
A connector's secrets — env values, a bearer token, an Authorization: header — are stored in plaintext in mcp-desktop.json on that machine. That is unavoidable: the adapter has to hand the real token to the server. Protect the file the way you protect ~/.aws/credentials. The masked input in /connect and the app is screen-share hygiene, not a storage claim.
Editing connectors from the phone or web app is the one place that's different, because the relay in between is untrusted: secrets are write-only in both directions. A listing returns env and header names and which of them are set, never a value; a value you type on the phone is sealed to that terminal's pinned key before it leaves the device, and the whole save is signed by your account — so the relay can neither read a credential nor forge a connector.
Hosted (Harbor) agents are OAuth-only by design: no stdio child processes and no stored tokens, because a hosted tenant's home is tmpfs and a durable secret would have to rest somewhere we could read. During the current preview they carry no connectors at all.
Full setup and troubleshooting live in the CLI guide.
Billing & limits
Usage is billed to your Privateer credit balance at your plan's rate — chat metered by the usage token counts on each completion, media metered per unit of output (see what media costs). Top up and see your balance in the app. Requests are subject to your plan's per-minute rate limit, daily message cap, and a small minimum-balance check — the same limits that govern the app.
| Limit | Default | On breach |
|---|---|---|
| Requests per minute | 120, over a rolling 60-second window (plans may raise this) | 429 RATE_LIMITED with retryAfter |
| Requests in flight at once | 8 — a pool of its own, separate from the app's and the agent's | 429 CONCURRENCY_LIMIT with cap |
| Daily messages | Your plan's cap | 429 naming the tier |
| Minimum balance | Checked before each billable call | 402 INSUFFICIENT_FUNDS with a top-up link |
| Audio upload size | 25 MB per file on /v1/audio/transcriptions | The upload is rejected before transcription |
Retry 429s with backoff, honouring retryAfter when present. A CONCURRENCY_LIMIT means slow down your fan-out, not your rate — waiting for an in-flight request to settle is what clears it.
No key, no charge you can't see. Every billed request writes a usage record (model, tokens, cost) you can reconcile — prompts and responses are never stored.
Errors
There are two error shapes, and it matters which one you get. Errors raised by an endpoint itself — bad input, a provider failure, a blocked model — use OpenAI's envelope, so an OpenAI SDK parses them normally. Errors raised by the gates that run before the endpoint (rate limit, credit check, plan caps, concurrency) return a flat body instead. An SDK that only reads error.message will see nothing on those; read the top-level code as well.
{ "error": { "message": "…", "type": "invalid_request_error", "code": "INVALID_REQUEST" } }
{ "code": "RATE_LIMITED", "message": "…", "retryAfter": 37 }
| Status | Code | Shape | Meaning |
|---|---|---|---|
400 | INVALID_REQUEST, PROMPT_REQUIRED, QUERY_REQUIRED, URL_REQUIRED, URL_INVALID | envelope | Malformed body (e.g. missing messages, empty input, no query, a relative or non-http URL) or an unknown model. |
401 | invalid_api_key | envelope | Missing, malformed, or revoked key; or the key's account isn't active. |
402 | INSUFFICIENT_FUNDS | flat or envelope | Balance too low. The pre-flight credit check is flat and adds balance, required and topUpUrl; a shortfall found mid-request comes back in the envelope. |
403 | ZDR_MEDIA_BLOCKED | envelope | The model has no Zero-Data-Retention endpoint and your account requires ZDR. See sound effects. |
403 | plan-gated feature | flat | Your plan doesn't include this (e.g. video). Adds effectiveTier, requiredTier, feature, graceEndsAt. |
422 | SPRITE_KEY_EMPTY | envelope | The chroma key removed the subject along with the backdrop, so there is nothing to pack. The request was well formed — lower chroma.tolerance or film against a cleaner backdrop. See sprite sheets. |
429 | RATE_LIMITED | flat | Over the per-minute limit. Adds retryAfter in seconds. |
429 | CONCURRENCY_LIMIT | flat | Too many requests in flight at once. Adds cap. |
429 | daily cap | flat | Plan's daily message cap reached. Adds effectiveTier and cap. |
502 | IMAGE_GEN_FAILED, INFERENCE_ERROR, SEARCH_FAILED, FETCH_FAILED | envelope | The provider failed. The upstream status and detail are folded into message so it's actionable. A failed search isn't billed. |
503 | SPRITE_GENERATION_UNAVAILABLE | envelope | This deployment has no video decoder, so frames cannot be read out of a rendered clip. Raised before anything is billed; POST /v1/sprites still packs frames you already have. |
503 | ZDR_KEY_UNAVAILABLE | envelope | ZDR routing was required and no compliant key was available. This fails loudly by design rather than quietly falling back. |
Provider errors are forwarded rather than reinterpreted, so an unfamiliar code in the envelope generally came from upstream.