← Back to overview
    Research & Development

    RTX PRO 4500 Blackwell: 3.4x LLM Throughput Through Runtime Optimization

    Same GPU, same main model, up to 3.4x the output: OPAIRS Runtime 3 raises the throughput of GPT-OSS-20B on the RTX PRO 4500 Blackwell to up to 2,637 tokens/s. At the same time, the tests show why Qwen3.8-27B will not take over the production stack for now.

    OPAIRS Runtime 3 on NVIDIA RTX PRO 4500 Blackwell with GPT-OSS-20B and up to 2,637 tokens per second
    OPAIRS Runtime 3: up to 3.4x output on the same RTX PRO 4500 Blackwell.

    The economics of local AI infrastructure do not depend on the purchase price of a GPU alone. What matters is the productive throughput that can be sustained from the available hardware under real load. For professional Blackwell cards in particular, this metric is gaining importance: when hardware capital costs rise, the price of the available inference capacity rises with them. OPAIRS therefore did not start by looking for a larger GPU, but examined the existing inference environment on the RTX PRO 4500 Blackwell. The focus was on the scheduler configuration, the parallelism that can actually be used, and the question of which model delivers the best balance of quality, performance, and availability on this hardware.

    The result of the tests: GPT-OSS-20B remains the main model of the OPAIRS stack for now. Through changes to the runtime and scheduler configuration, the measured output throughput on a single RTX PRO 4500 Blackwell was increased from around 766 to up to 2,637 tokens/s. That is an increase by a factor of 3.4, without replacing the GPU or the production model.

    Qwen3.8-27B in Comparison: Newer Model Generation, Lower Throughput

    Benchmark comparison of GPT-OSS-20B and Qwen3.8-27B on RTX PRO 4500 Blackwell
    Single-GPU benchmark on RTX PRO 4500 Blackwell: in this load profile, GPT-OSS-20B achieves more than twice the output throughput.

    A Newer Architecture Does Not Automatically Mean Higher Production Performance

    Qwen3.8-27B-NVFP4 was evaluated as a possible successor to GPT-OSS-20B on the same RTX PRO 4500 Blackwell. The model could be run stably with a 32K context. Under parallel load, however, a clear difference emerged: in the comparison, GPT-OSS-20B achieved 766.47 output tokens/s with a mean TTFT of 6.97 seconds and a mean TPOT of 12.78 ms. Qwen3.8-27B achieved 316.61 output tokens/s with a mean TTFT of 13.11 seconds and a mean TPOT of 36.15 ms. In this test scenario, Qwen thus reaches around 41 percent of the output throughput of GPT-OSS-20B. For the current single-GPU configuration, GPT-OSS-20B therefore continues to deliver the better balance of performance and availability.

    The First Bottleneck Was Not the GPU but the Scheduler

    The previous production configuration of GPT-OSS-20B used --max-num-seqs 10. This setting had provided a stable starting point for operation, but it limited possible parallelization more than necessary. As the number of requests increased, the load analysis showed a clear pattern: time-to-first-token rose significantly, while overall throughput remained in the range of roughly 760 to 770 tokens/s. The available GPU capacity was therefore not consistently translated into additional output.

    OPAIRS Runtime 3 vLLM max-num-seqs benchmark from 10 to 60 parallel sequences
    Concurrency sweep of the runtime: the scheduler was the first limiting factor, not the available GPU compute.

    From 10 to 60 Concurrently Managed Sequences

    OPAIRS therefore systematically examined max-num-seqs across different request, prompt, and output profiles. The new Runtime 3 configuration sets --max-num-seqs 60. The parameter does not statically reserve resources for 60 users. It defines the maximum number of sequences managed concurrently by the vLLM scheduler. The change removes the previously observed scheduler cap and makes it possible to utilize the existing GPU far better under parallel load.

    Up to 60 Parallel Sequences Does Not Mean 60 Users

    OPAIRS Runtime 3 with parallel users, APIs, agents, tool calls, and RAG processes
    A session is not the same as a request: agents and tool calls generate additional parallel inference load.

    Clients, Tool Calls, and Agents Share the Same Runtime

    In an industrial AI system, the number of human users is only part of the actual inference load. A single user can trigger multiple tool calls and further model calls through an agent. At the same time, APIs, retrieval processes, background jobs, and automated workflows access the same inference instance. The relevant technical metric is therefore not 60 users, but up to 60 inference sequences managed concurrently by the scheduler on a single GPU. These can originate from user interactions, API clients, agents, tool calls, RAG processes, or automated systems.

    The Result: From 766 to Up to 2,637 Tokens/s

    OPAIRS Runtime 3 benchmark result: 766 to 2,637 tokens per second
    Same GPU, same main model: up to 2,637 output tokens/s through runtime and scheduler optimization.

    3.4x Output on the Same RTX PRO 4500 Blackwell

    The optimization changes neither the GPU used nor the main model. It changes how the available compute capacity is used under parallel load. With GPT-OSS-20B on the RTX PRO 4500 Blackwell, the measured output rises from a baseline of around 766 tokens/s to a peak of up to 2,637 tokens/s. The peak value is not the only thing that matters here. The configuration was tested across short, medium, long, and very long load profiles. Across more than 15 test series, no failed requests were observed. Only after these load tests was max-num-seqs 60 adopted as the new production configuration.

    Performance and Availability Are Only Two Parts of the Architecture

    For industrial AI, high throughput alone is not enough. Alongside performance and availability, the domain-specific quality must be right. OPAIRS therefore deliberately does not follow the approach of solving every task with a single, as-large-as-possible model. GPT-OSS-20B remains the central main model for general requests and at the same time handles orchestration within the platform. When the platform recognizes a specialized industrial task, the corresponding workload can be handed off in a targeted way to a smaller, fine-tuned model on a different GPU.

    OPAIRS multi-GPU architecture with GPT-OSS-20B as the main model and specialized fine-tuned models
    The right model for the right task: GPT-OSS-20B orchestrates specialized models on separate GPU resources.

    GPT-OSS-20B as the Main Model, Specialized Models as Domain Instances

    GPT-OSS-20B handles the central interaction and decides, together with the OPAIRS platform, when additional specialization is required. For tasks in areas such as production, maintenance, or SQL and ABAP, smaller fine-tuned models are available. These run on separate GPUs. This not only applies domain knowledge in a targeted way, but also separates the load: a complex SQL query or an industry-specific maintenance analysis does not have to permanently occupy compute capacity on the main model GPU. After the specialized processing, the result flows back into the central OPAIRS context.

    Training on Leonardo, Inference on the OPAIRS Infrastructure

    EuroHPC Leonardo training of specialized models and local multi-GPU deployment with OPAIRS
    Training on European HPC infrastructure, production inference locally within the OPAIRS platform.

    Specialization Begins Before Local Deployment

    The specialized models are created as part of the OPAIRS training program on the Leonardo BOOSTER in Bologna. Through EuroHPC, OPAIRS has 5,000 GPU hours available for this purpose. The compute-intensive training and evaluation phase is thus separated from later production inference. After fine-tuning, the models are integrated into the OPAIRS platform and operated locally on the respective designated GPU resources. Current internal evaluations already show measurable quality differences between specialized models and base models not trained on the respective domain. The full results of these evaluations, as well as the multi-model orchestration, will be presented in a dedicated follow-up article.

    More about the training program on Leonardo: https://opairs-systems.com/insights/eurohpc-spezialisierte-ki-agenten-fertigung

    Even 2,637 Tokens/s May Not Be the End

    RTX PRO Blackwell SM120 Marlin fallback and native MXFP4 CUTLASS FlashInfer kernel
    Next area of investigation: from the stable Marlin fallback to native SM120 MXFP4 kernels.

    The Next Optimization Step: Native MXFP4 on SM120

    On the RTX PRO Blackwell architecture, GPT-OSS-20B currently still uses a Marlin path for MXFP4. Native support for SM120 workstation GPUs remains the subject of active development in the open-source inference frameworks involved. OPAIRS is therefore currently investigating whether today's Marlin fallback can be replaced by a stable native SM120 MXFP4 path based on the available CUTLASS and FlashInfer work. Published comparison measurements on other SM120 hardware show improvements over Marlin in the range of roughly 13 to 28 percent, depending on the load profile. For OPAIRS, this results in a technical target of around 20 percent of additional potential. This value is explicitly not yet a measured benchmark on the RTX PRO 4500 and will be validated separately before adoption into the production runtime.

    Technical References on SM120 and MXFP4 Support

    • vLLM #31085 – Native NVFP4/MXFP4 support and backend selection for SM120: https://github.com/vllm-project/vllm/issues/31085
    • vLLM #30135 – Marlin fallback on RTX PRO Blackwell / SM120: https://github.com/vllm-project/vllm/issues/30135
    • vLLM #33416 – SM120 capability checks in the quantization and kernel path: https://github.com/vllm-project/vllm/issues/33416
    • FlashInfer #2847 – MXFP4 MoE on SM120 and comparison measurements against Marlin: https://github.com/flashinfer-ai/flashinfer/issues/2847
    • NVIDIA CUTLASS #3096 – SM120 FP4/MXFP4 kernel and performance investigations: https://github.com/NVIDIA/cutlass/issues/3096

    Conclusion: Optimization Before Hardware Scaling

    The tests show that with local LLM inference, considerable performance reserves can lie outside the model itself. On a single RTX PRO 4500 Blackwell, OPAIRS was able to increase the output throughput of GPT-OSS-20B from around 766 to up to 2,637 tokens/s through runtime and scheduler optimization. GPT-OSS-20B thus remains the main model of the system. Qwen3.8-27B continues to be evaluated, but on the tested single-GPU configuration it currently does not reach the required balance of throughput and latency.

    For industry-specific tasks, OPAIRS is in parallel relying on smaller models trained on Leonardo and fine-tuned for defined domains. GPT-OSS-20B remains the central instance for general processing and orchestration, while specialized workloads are offloaded in a targeted way to other models and GPU resources. Performance, availability, and domain-specific quality are thus not treated as properties of a single model, but as the result of the entire system architecture.

    At the same time, the runtime work is not yet complete. OPAIRS is currently investigating whether the existing Marlin fallback can be replaced by a stable native SM120 MXFP4 path. Published results on comparable SM120 hardware point to further performance potential. The next optimization stage therefore does not necessarily consist of a larger GPU or a larger model, but of the combination of runtime optimization, load separation, model specialization, and the fullest possible use of the existing Blackwell hardware.

    The results of our current evaluations of the specialized SQL, production, and maintenance models, as well as how they interact with GPT-OSS-20B, will be published in a follow-up article.

    For technical collaboration, benchmark results, or joint industrial projects: office@opairs-systems.com

    More insights

    OPAIRS SQL Agent benchmark: pass rate of all six LoRA adapters for PostgreSQL, T-SQL and Apache Iceberg compared with Claude Opus 4.8, GPT-OSS-20B and the untuned base models
    Research & Development

    OPAIRS SQL Agent: Comparing Six LoRA Adapters for Industrial Databases

    Six LoRA adapters, a 26-question catalog spanning PostgreSQL, T-SQL and Apache Iceberg databases, two external reference models: OPAIRS has systematically evaluated its SQL agent. Two Granite adapters lead the field. For production use, however, the deciding factors are not only answer quality but also speed, memory footprint, concurrency and the available context from ERP, MES, PLM and other industrial systems.

    Read article
    NVIDIA Inception Program badge – OPAIRS Systems is a program member
    Partnership

    OPAIRS Joins the NVIDIA Inception Program

    GPU-accelerated LLM inference on-premise: what NVIDIA Inception means for the OPAIRS stack, and why sovereign AI and high-performance hardware must go hand in hand.

    Read article