Red Hat AI 3.5 adds safety evaluation, agent controls and shared-GPU scheduling
The release moves model checks, inference observability and multi-tenant resource controls into the platform as Red Hat targets production AI operations.
Red Hat has released Red Hat AI 3.5, packaging model evaluation, inference observability, multi-tenant GPU controls and agent-development components into the company’s enterprise AI platform. The release is generally available, including through Red Hat AI Factory with NVIDIA, according to the company announcement.
Safety checks move before deployment
The clearest operational change is EvalHub, now generally available. Red Hat says it can automate safety benchmarking and produce auditable compliance reports for custom models, retrieval-augmented generation systems and agents. Models in the platform catalog can also carry Garak results covering safety, personally identifiable information exposure and toxicity risk.
That does not make model safety automatic, but it gives platform teams a common place to collect evidence before deployment rather than treating evaluation as a separate project. Red Hat says selected catalog models are also marked as validated for tool calling, an important distinction as agent systems begin invoking external services.
More control over shared inference
For shared GPU infrastructure, the release adds fair-share scheduling and priority-aware serving. The latter applies admission control and priority-based request routing so latency-sensitive inference can take precedence while background work consumes spare capacity. Controlled model rollout is intended to steer traffic during model updates.
Red Hat AI 3.5 also officially supports hosted control planes on OpenShift Virtualization. The design gives tenants separate control planes while consolidating hardware, with virtual machines providing an additional isolation boundary for AI workloads. CPU offloading is generally available, while storage offloading remains a developer preview.
The OpenShift AI 3.5 documentation now points to the 3.5.1 release notes, indicating that operators should check the maintained documentation for current fixes and known issues rather than relying only on the launch summary.
Agents and observability
The platform now generally supports the Responses API and built-in RAG, with NeMo Guardrails integration intended to intercept malicious tool calls. AutoRAG adds pgvector support, multilingual documents, conversational testing and a visual pipeline, while AI Hub introduces templates for code review, document processing and research workflows.
For operations teams, new dashboards expose model and agent performance, GPU utilization and per-user token consumption. MLflow tracing adds visibility into agent execution, and Red Hat says showback data can be accessed by non-administrators.
The release is notable less for a single new model-serving feature than for bringing evaluation, scheduling, isolation and usage accounting into one supported platform. Teams adopting it should still validate the maturity level of individual components: several additions are previews or early access, including storage offloading, the Kubeflow Spark Operator integration and multimodal serving through vLLM Omni.
sources
comments · 0