Red Hat packages NVIDIA AI-Q as an OpenShift AI research-agent quickstart
The deployment guide combines multi-agent routing, vLLM model serving and an optional observability stack, with local-GPU and NVIDIA-hosted paths.
Red Hat has published a hands-on quickstart for deploying a multi-agent research application on Red Hat AI Factory with NVIDIA. The guide adapts NVIDIA’s AI-Q Blueprint for Red Hat AI environments and adds deployment choices for locally served models or NVIDIA-hosted inference.
What the quickstart assembles
The application routes requests between simple responses, shallow tool-assisted research and deeper investigations that use planning, subagents and report generation. It can draw from web and academic search, uploaded files and an optional retrieval-augmented generation layer, while preserving citations in the output, according to the deployment documentation.
The underlying pattern combines NVIDIA NeMo Agent Toolkit with LangChain Deep Agents. For a local deployment, the guide serves three model roles through vLLM and KServe: RedHatAI/gpt-oss-120b as orchestrator, RedHatAI/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 for intent routing and research, and nvidia/Nemotron-Mini-4B-Instruct for summarization. Red Hat says the documented configuration was tested with OpenShift 4.20 and OpenShift AI 3.3.2.
Red Hat’s accompanying demonstration shows the workflow selecting a research path, collecting evidence and producing a report. It also shows operators tracking execution with OpenTelemetry and MLflow, watching model-serving metrics in Grafana and running evaluations from OpenShift AI workbenches.
Two deployment paths, with different costs
The local-vLLM option keeps model serving and data inside the cluster, but its standard configuration calls for three 80GB NVIDIA H100 or A100 GPUs. The guide describes a two-H100 minimum using Multi-Instance GPU partitioning. Teams can instead use NVIDIA NGC-hosted models without local GPUs, trading infrastructure control for a cloud API and pay-per-use inference.
The optional observability deployment installs OpenShift Logging with LokiStack, Grafana, OpenTelemetry Collector, user-workload monitoring and MLflow. The default application deployment uses a 10GB PostgreSQL persistent volume, while ChromaDB and application data use ephemeral storage unless an administrator adds another persistent volume.
The operational caveat
This is an engineering quickstart, not a blanket support statement. Red Hat explicitly says the material is authored by its experts but has not been tested on every supported configuration. The repository also uses prebuilt images based on NVIDIA AI-Q 2.1.0 with Red Hat-specific patches, and the documentation exposes those deployment files and patches for inspection.
For platform teams, the useful development is the amount of the agent stack that is made concrete: model endpoints, routing, retrieval, traces, metrics, evaluation and cleanup are all represented in deployable manifests. The caveat is equally concrete: the local path has a substantial GPU floor, and the faster hosted path sends inference to NVIDIA’s service.
sources
comments · 0