Fair Use Policy
The restrictions below are current fair usage limitations for each of the managed LLM models, in order to make the LLM usage experience stable for everyone.
The LLM gateway enforces one limit automatically: 200,000 output tokens per minute per API token and model, beyond which requests receive HTTP 429. Input tokens are not counted. Every response carries x-ratelimit-remaining and x-ratelimit-reset (seconds until the minute rolls over) headers, so scripts can back off before being rejected. The concurrency limits below are not enforced automatically. Please report to the Nautilus Artificial Intelligence/Machine Learning channel in Natilus Support if you observe lagging requests and high request volume in Grafana.
| Maximum Per-User Concurrency | Models |
|---|---|
2 | kimi, glm-5, deepseek-v4-flash |
8 | minimax-m2, qwen3-small, gemma, gemma-small |
16 | qwen3, gpt-oss, qwen3-embedding |
San Diego Supercomputer Center (SDSC) and Internet2 have contributed their GPU nodes for managed LLM inference. Therefore, users affiliated with these organizations are granted twice the limits and separately arrangeable higher-volume sessions.
The National Research Platform is non-profit and non-commercial. All usage of the cluster, including LLMs, is non-profit and non-commercial, as per the AUP.

This work was supported in part by National Science Foundation (NSF) awards CNS-1730158, ACI-1540112, ACI-1541349, OAC-1826967, OAC-2112167, CNS-2100237, CNS-2120019.