Fair Use Policy
The restrictions below are current fair usage limitations for each of the managed LLM models, in order to make the LLM usage experience stable for everyone.
The below limits are currently not automatically enforced, but automatic rate limiting in the LLM gateway is coming soon. Please report to the Nautilus Artificial Intelligence/Machine Learning channel in Natilus Support if you observe lagging requests and high request volume in Grafana.
| Maximum Per-User Concurrency | Models |
|---|---|
2 | kimi, glm-5 |
8 | minimax-m2, qwen3-small, gemma, gemma-small |
16 | qwen3, gpt-oss, qwen3-embedding |
San Diego Supercomputer Center (SDSC) and Internet2 have contributed their GPU nodes for managed LLM inference. Therefore, users affiliated with these organizations are granted twice the limits and separately arrangeable higher-volume sessions.
The National Research Platform is non-profit and non-commercial. All usage of the cluster, including LLMs, is non-profit and non-commercial, as per the AUP.
