GPT-OSS 120B
- Input / 1M tokens
- $0.13875
- Output / 1M tokens
- $0.555
- Availability
- Non-callable
AI CLOUD · GPU INFRASTRUCTURE
Explore open-model services, GPU hosting and partner capacity. Choose how you build—with clearer ownership, deployment and cost boundaries.
Models. Compute. Ownership.
Application layer
Compute layer
Infrastructure layer
BUILD YOUR WAY
Start with your workload and ownership model. Compare the requirements before making an infrastructure commitment.
Compare model profiles, API contracts and price units for your application.
Explore inference 02 / AI COLOCATIONPlan power, cooling, connectivity and operations for customer-owned servers.
Explore GPU hosting 03 / PARTNER CAPACITYUnderstand supply qualification, ownership and operating responsibilities.
Explore partnershipsDEVELOPER EXPERIENCE
Start with a familiar API contract. PacinfraX-specific controls stay explicit, versioned, and optional.
# Contract example; public endpoint is not enabled.
# Set PACINFRAX_API_KEY securely in your environment.
curl --no-buffer --fail-with-body --max-time 20 \
--proto '=https' --tlsv1.2 \
https://api.pacinfrax.ai/v1/chat/completions \
-H "Authorization: Bearer $PACINFRAX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model-id>",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'
# Prints raw server-sent events, not parsed assistant text.# Python 3 stdlib; contract endpoint is not enabled.
import json
import os
import sys
import urllib.error
import urllib.request
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, req, fp, code, msg, headers, newurl):
return None
key = os.environ.get("PACINFRAX_API_KEY")
if not key:
raise SystemExit("Set PACINFRAX_API_KEY securely first.")
body = {
"model": "<model-id>",
"messages": [{"role": "user", "content": "Hello"}],
"stream": True,
}
request = urllib.request.Request(
"https://api.pacinfrax.ai/v1/chat/completions",
data=json.dumps(body).encode("utf-8"),
headers={"Authorization": "Bearer " + key,
"Content-Type": "application/json"},
method="POST",
)
try:
with urllib.request.build_opener(NoRedirect).open(request, timeout=20) as response:
if response.headers.get_content_type() != "text/event-stream":
raise SystemExit("Expected a streaming response.")
# Write raw SSE bytes; do not decode partial UTF-8 chunks.
while chunk := response.read1(4096):
sys.stdout.buffer.write(chunk)
sys.stdout.buffer.flush()
except urllib.error.HTTPError as error:
raise SystemExit("HTTP request rejected: " + str(error.code)) from None
except (urllib.error.URLError, OSError):
raise SystemExit("Streaming request failed.") from None// Node.js 20+ built-in fetch; contract endpoint is not enabled.
const key = process.env.PACINFRAX_API_KEY;
if (!key) throw new Error("Set PACINFRAX_API_KEY securely first.");
async function streamExample(): Promise<void> {
const response = await fetch(
"https://api.pacinfrax.ai/v1/chat/completions", {
method: "POST",
redirect: "error",
signal: AbortSignal.timeout(20_000),
headers: {
Authorization: "Bearer " + key,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "<model-id>",
messages: [{ role: "user", content: "Hello" }],
stream: true,
}),
},
);
if (!response.ok || !response.body ||
response.headers.get("content-type")?.split(";")[0].trim() !== "text/event-stream") {
await response.body?.cancel();
throw new Error("Expected a successful streaming response.");
}
const reader = response.body.getReader();
const decoder = new TextDecoder();
try {
// Print raw SSE frames, not parsed assistant text.
while (true) {
const { value, done } = await reader.read();
if (done) break;
process.stdout.write(decoder.decode(value, { stream: true }));
}
process.stdout.write(decoder.decode());
} finally {
await reader.cancel().catch(() => {});
reader.releaseLock();
}
}
streamExample().catch(() => {
console.error("Streaming request failed.");
process.exitCode = 1;
});Contract examples only: the public API endpoint is not enabled. Replace <model-id> with an authorized callable model ID before use. Examples print raw SSE events; they do not invoke a model on this page.
MODEL CATALOG
Compare reference profiles and indicative token prices. Open a model passport for license, runtime and deployment requirements.
Indicative prices, not billable rates. Partial draft—not current capacity. License, runtime, region and commercial acceptance remain pending. Reference checked 2026-10-03. Explore all models and service units →
MODEL & SERVICE PRICING
Compare model and developer-service estimates separately from GPU hosting. Understand the unit before choosing your deployment.
Indicative estimates only—not billable rates or a quote. Listed services are not available for invocation. Reference checked 2026-10-03.
TRUST BY DESIGN
Understand project boundaries, credential handling and diagnostic practices before choosing your deployment.
Review security boundariesExpiry and revocation are part of the credential contract. Secrets are designed to be shown once.
Operational analytics use sanitized request metadata; prompt and output sharing stays opt-in.
Review required regions and data handling before committing to a deployment.
YOUR NEXT BUILD
Prepare a workload brief locally before any capacity, region, price or timeline commitment. Customer intake is not enabled; nothing is submitted automatically.