PACINFRAX · PRODUCT

Choose how your model serves requests.

Compare shared requests, dedicated endpoints and asynchronous batches.

Your decision

Match the serving mode to latency, demand and operating responsibility.

  • Shared API: variable request demand with model-level input and output units. Compare model revision and compatibility before integration.
  • Dedicated endpoint: a controlled serving configuration with an accepted isolation, scaling and drain policy.
  • Batch: queued work with a completion window, result retention and cancellation policy rather than an interactive latency promise.

Public service activation is not available here. Compare requirements before choosing a deployment; no resource, price or capacity is reserved.

Compare model references

Continue your decision