Dashboard / Services / Inference API
Inference API
Requests today
2.4M
12% above the daily average
Success rate
99.98%
Within the 99.9% target
P95 latency
182 ms
18 ms below the alert threshold
Service details
Current production configuration.
- Region
- us-central-1
- Model
- Llama 3.3 70B Instruct
- Deployment
- inference-api-v42
Platform administrators
12 members can deploy, configure, and manage access.
ML operations
28 members can deploy and update service configuration.