live wire
▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel▸JAVA · Quarkus 4.0.0.Beta1 moves to Java 21, adds HTTP/3 and starts extension migration (Oct. 1)Quarkus▸SECURITY · X41 shows shared /dev/shm can turn Envoy hot restart into cross-container lateral movementX41 D-Sec▸DATA · AWS and Red Hat map Confluent Platform on ROSA with HCP, CFK and OpenShift security controlsAWS IBM & Red Hat▸API · Red Hat resolves intermittent 3scale API Manager latencyRed Hat Status▸AI · IBM shows Maximo workflows exposed as approval-gated MCP tools on OpenShiftIBM Community▸AI · vLLM adds day-zero NVIDIA Vera Rubin support and reports 7.8× per-GPU throughputvLLM▸INTEGRATION · Apache Camel 4.23 makes Kamelets visible to AI tooling and validationApache Camel▸SECURITY · OpenShift 4.14.75 fixes five CVEs, including two SQLite code-execution flawsRed Hat Customer Portal▸SUPPLY CHAIN · Red Hat maps CRA-ready open source practices as EU reporting rules take effectRed Hat Blog▸AI · Red Hat AI Inference on IBM Cloud adds an OpenAI-compatible Embeddings APIIBM Cloud▸API · Red Hat investigates degraded 3scale API Management SaaS APIsRed Hat Status▸PLATFORM · Red Hat and Cloudera validate a 100-VM analytics stack on OpenShift VirtualizationRed Hat Blog▸DEVELOPER HUB · Red Hat maps a four-zone, quota-aware Dev Spaces architectureRed Hat Developer▸INTEGRATION · Camel 4.23 teaches agent tools to discover and validate KameletsApache Camel
upstreambeat.ai
releaseAI

Red Hat AI 3.5 adds safety evaluation, agent controls and shared-GPU scheduling

The release moves model checks, inference observability and multi-tenant resource controls into the platform as Red Hat targets production AI operations.

Red Hat AI 3.5 moves safety checks and GPU controls into one platform.
AI-generated illustration
By The News Desk· Sep 16, 2026the quick take — two AI hosts go live when you do

Red Hat has released Red Hat AI 3.5, packaging model evaluation, inference observability, multi-tenant GPU controls and agent-development components into the company’s enterprise AI platform. The release is generally available, including through Red Hat AI Factory with NVIDIA, according to the company announcement.

Safety checks move before deployment

The clearest operational change is EvalHub, now generally available. Red Hat says it can automate safety benchmarking and produce auditable compliance reports for custom models, retrieval-augmented generation systems and agents. Models in the platform catalog can also carry Garak results covering safety, personally identifiable information exposure and toxicity risk.

That does not make model safety automatic, but it gives platform teams a common place to collect evidence before deployment rather than treating evaluation as a separate project. Red Hat says selected catalog models are also marked as validated for tool calling, an important distinction as agent systems begin invoking external services.

More control over shared inference

For shared GPU infrastructure, the release adds fair-share scheduling and priority-aware serving. The latter applies admission control and priority-based request routing so latency-sensitive inference can take precedence while background work consumes spare capacity. Controlled model rollout is intended to steer traffic during model updates.

Red Hat AI 3.5 also officially supports hosted control planes on OpenShift Virtualization. The design gives tenants separate control planes while consolidating hardware, with virtual machines providing an additional isolation boundary for AI workloads. CPU offloading is generally available, while storage offloading remains a developer preview.

The OpenShift AI 3.5 documentation now points to the 3.5.1 release notes, indicating that operators should check the maintained documentation for current fixes and known issues rather than relying only on the launch summary.

Agents and observability

The platform now generally supports the Responses API and built-in RAG, with NeMo Guardrails integration intended to intercept malicious tool calls. AutoRAG adds pgvector support, multilingual documents, conversational testing and a visual pipeline, while AI Hub introduces templates for code review, document processing and research workflows.

For operations teams, new dashboards expose model and agent performance, GPU utilization and per-user token consumption. MLflow tracing adds visibility into agent execution, and Red Hat says showback data can be accessed by non-administrators.

The release is notable less for a single new model-serving feature than for bringing evaluation, scheduling, isolation and usage accounting into one supported platform. Teams adopting it should still validate the maturity level of individual components: several additions are previews or early access, including storage offloading, the Kubeflow Spark Operator integration and multimodal serving through vLLM Omni.

Filed by The News Desk. Corrections: desk@upstreambeat.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.