RTX PRO 4500 Blackwell: 3.4x LLM Throughput Through Runtime Optimization
Same GPU, same main model, up to 3.4x the output: OPAIRS Runtime 3 raises the throughput of GPT-OSS-20B on the RTX PRO 4500 Blackwell to up to 2,637 tokens/s. At the same time, the tests show why Qwen3.8-27B will not take over the production stack for now.

The economics of local AI infrastructure do not depend on the purchase price of a GPU alone. What matters is the productive throughput that can be sustained from the available hardware under real load. For professional Blackwell cards in particular, this metric is gaining importance: when hardware capital costs rise, the price of the available inference capacity rises with them. OPAIRS therefore did not start by looking for a larger GPU, but examined the existing inference environment on the RTX PRO 4500 Blackwell. The focus was on the scheduler configuration, the parallelism that can actually be used, and the question of which model delivers the best balance of quality, performance, and availability on this hardware.
The result of the tests: GPT-OSS-20B remains the main model of the OPAIRS stack for now. Through changes to the runtime and scheduler configuration, the measured output throughput on a single RTX PRO 4500 Blackwell was increased from around 766 to up to 2,637 tokens/s. That is an increase by a factor of 3.4, without replacing the GPU or the production model.
Qwen3.8-27B in Comparison: Newer Model Generation, Lower Throughput

A Newer Architecture Does Not Automatically Mean Higher Production Performance
Qwen3.8-27B-NVFP4 was evaluated as a possible successor to GPT-OSS-20B on the same RTX PRO 4500 Blackwell. The model could be run stably with a 32K context. Under parallel load, however, a clear difference emerged: in the comparison, GPT-OSS-20B achieved 766.47 output tokens/s with a mean TTFT of 6.97 seconds and a mean TPOT of 12.78 ms. Qwen3.8-27B achieved 316.61 output tokens/s with a mean TTFT of 13.11 seconds and a mean TPOT of 36.15 ms. In this test scenario, Qwen thus reaches around 41 percent of the output throughput of GPT-OSS-20B. For the current single-GPU configuration, GPT-OSS-20B therefore continues to deliver the better balance of performance and availability.
The First Bottleneck Was Not the GPU but the Scheduler
The previous production configuration of GPT-OSS-20B used --max-num-seqs 10. This setting had provided a stable starting point for operation, but it limited possible parallelization more than necessary. As the number of requests increased, the load analysis showed a clear pattern: time-to-first-token rose significantly, while overall throughput remained in the range of roughly 760 to 770 tokens/s. The available GPU capacity was therefore not consistently translated into additional output.

From 10 to 60 Concurrently Managed Sequences
OPAIRS therefore systematically examined max-num-seqs across different request, prompt, and output profiles. The new Runtime 3 configuration sets --max-num-seqs 60. The parameter does not statically reserve resources for 60 users. It defines the maximum number of sequences managed concurrently by the vLLM scheduler. The change removes the previously observed scheduler cap and makes it possible to utilize the existing GPU far better under parallel load.
Up to 60 Parallel Sequences Does Not Mean 60 Users

Clients, Tool Calls, and Agents Share the Same Runtime
In an industrial AI system, the number of human users is only part of the actual inference load. A single user can trigger multiple tool calls and further model calls through an agent. At the same time, APIs, retrieval processes, background jobs, and automated workflows access the same inference instance. The relevant technical metric is therefore not 60 users, but up to 60 inference sequences managed concurrently by the scheduler on a single GPU. These can originate from user interactions, API clients, agents, tool calls, RAG processes, or automated systems.
The Result: From 766 to Up to 2,637 Tokens/s

3.4x Output on the Same RTX PRO 4500 Blackwell
The optimization changes neither the GPU used nor the main model. It changes how the available compute capacity is used under parallel load. With GPT-OSS-20B on the RTX PRO 4500 Blackwell, the measured output rises from a baseline of around 766 tokens/s to a peak of up to 2,637 tokens/s. The peak value is not the only thing that matters here. The configuration was tested across short, medium, long, and very long load profiles. Across more than 15 test series, no failed requests were observed. Only after these load tests was max-num-seqs 60 adopted as the new production configuration.
Performance and Availability Are Only Two Parts of the Architecture
For industrial AI, high throughput alone is not enough. Alongside performance and availability, the domain-specific quality must be right. OPAIRS therefore deliberately does not follow the approach of solving every task with a single, as-large-as-possible model. GPT-OSS-20B remains the central main model for general requests and at the same time handles orchestration within the platform. When the platform recognizes a specialized industrial task, the corresponding workload can be handed off in a targeted way to a smaller, fine-tuned model on a different GPU.

GPT-OSS-20B as the Main Model, Specialized Models as Domain Instances
GPT-OSS-20B handles the central interaction and decides, together with the OPAIRS platform, when additional specialization is required. For tasks in areas such as production, maintenance, or SQL and ABAP, smaller fine-tuned models are available. These run on separate GPUs. This not only applies domain knowledge in a targeted way, but also separates the load: a complex SQL query or an industry-specific maintenance analysis does not have to permanently occupy compute capacity on the main model GPU. After the specialized processing, the result flows back into the central OPAIRS context.
Training on Leonardo, Inference on the OPAIRS Infrastructure

Specialization Begins Before Local Deployment
The specialized models are created as part of the OPAIRS training program on the Leonardo BOOSTER in Bologna. Through EuroHPC, OPAIRS has 5,000 GPU hours available for this purpose. The compute-intensive training and evaluation phase is thus separated from later production inference. After fine-tuning, the models are integrated into the OPAIRS platform and operated locally on the respective designated GPU resources. Current internal evaluations already show measurable quality differences between specialized models and base models not trained on the respective domain. The full results of these evaluations, as well as the multi-model orchestration, will be presented in a dedicated follow-up article.
More about the training program on Leonardo: https://opairs-systems.com/insights/eurohpc-spezialisierte-ki-agenten-fertigung
Even 2,637 Tokens/s May Not Be the End

The Next Optimization Step: Native MXFP4 on SM120
On the RTX PRO Blackwell architecture, GPT-OSS-20B currently still uses a Marlin path for MXFP4. Native support for SM120 workstation GPUs remains the subject of active development in the open-source inference frameworks involved. OPAIRS is therefore currently investigating whether today's Marlin fallback can be replaced by a stable native SM120 MXFP4 path based on the available CUTLASS and FlashInfer work. Published comparison measurements on other SM120 hardware show improvements over Marlin in the range of roughly 13 to 28 percent, depending on the load profile. For OPAIRS, this results in a technical target of around 20 percent of additional potential. This value is explicitly not yet a measured benchmark on the RTX PRO 4500 and will be validated separately before adoption into the production runtime.
Technical References on SM120 and MXFP4 Support
- vLLM #31085 – Native NVFP4/MXFP4 support and backend selection for SM120: https://github.com/vllm-project/vllm/issues/31085
- vLLM #30135 – Marlin fallback on RTX PRO Blackwell / SM120: https://github.com/vllm-project/vllm/issues/30135
- vLLM #33416 – SM120 capability checks in the quantization and kernel path: https://github.com/vllm-project/vllm/issues/33416
- FlashInfer #2847 – MXFP4 MoE on SM120 and comparison measurements against Marlin: https://github.com/flashinfer-ai/flashinfer/issues/2847
- NVIDIA CUTLASS #3096 – SM120 FP4/MXFP4 kernel and performance investigations: https://github.com/NVIDIA/cutlass/issues/3096
Conclusion: Optimization Before Hardware Scaling
The tests show that with local LLM inference, considerable performance reserves can lie outside the model itself. On a single RTX PRO 4500 Blackwell, OPAIRS was able to increase the output throughput of GPT-OSS-20B from around 766 to up to 2,637 tokens/s through runtime and scheduler optimization. GPT-OSS-20B thus remains the main model of the system. Qwen3.8-27B continues to be evaluated, but on the tested single-GPU configuration it currently does not reach the required balance of throughput and latency.
For industry-specific tasks, OPAIRS is in parallel relying on smaller models trained on Leonardo and fine-tuned for defined domains. GPT-OSS-20B remains the central instance for general processing and orchestration, while specialized workloads are offloaded in a targeted way to other models and GPU resources. Performance, availability, and domain-specific quality are thus not treated as properties of a single model, but as the result of the entire system architecture.
At the same time, the runtime work is not yet complete. OPAIRS is currently investigating whether the existing Marlin fallback can be replaced by a stable native SM120 MXFP4 path. Published results on comparable SM120 hardware point to further performance potential. The next optimization stage therefore does not necessarily consist of a larger GPU or a larger model, but of the combination of runtime optimization, load separation, model specialization, and the fullest possible use of the existing Blackwell hardware.
The results of our current evaluations of the specialized SQL, production, and maintenance models, as well as how they interact with GPT-OSS-20B, will be published in a follow-up article.
For technical collaboration, benchmark results, or joint industrial projects: office@opairs-systems.com
More insights

OPAIRS SQL Agent: Comparing Six LoRA Adapters for Industrial Databases
Six LoRA adapters, a 26-question catalog spanning PostgreSQL, T-SQL and Apache Iceberg databases, two external reference models: OPAIRS has systematically evaluated its SQL agent. Two Granite adapters lead the field. For production use, however, the deciding factors are not only answer quality but also speed, memory footprint, concurrency and the available context from ERP, MES, PLM and other industrial systems.
Read articleOPAIRS Joins the NVIDIA Inception Program
GPU-accelerated LLM inference on-premise: what NVIDIA Inception means for the OPAIRS stack, and why sovereign AI and high-performance hardware must go hand in hand.
Read article