Red Hat quickstart turns embedding selection into an OpenShift AI pipeline
The proof of concept benchmarks embedding models, promotes the winner and deploys a CPU-based semantic-search service through KServe.
A new Red Hat AI Quickstarts reference build packages embedding-model evaluation, selection and deployment into one workflow for Red Hat OpenShift AI. The proof of concept uses ZenML to benchmark candidate models against technical-support queries, selects the strongest result, rebuilds the retrieval index and deploys a searchable service through KServe.
The project is aimed at a practical problem in retrieval-augmented generation: the quality of the context supplied to an application or agent depends heavily on the embedding model, but model selection is often detached from deployment. This build makes that choice a repeatable pipeline step rather than a one-time experiment.
What the pipeline does
The documented workflow prepares a TechQA retrieval benchmark, evaluates configured embedding candidates in parallel and records results in MLflow. It then re-encodes the corpus with the selected model, stores a versioned FAISS bundle in MinIO and creates or updates a CPU-based KServe InferenceService.
The deployed FastAPI application exposes a browser search interface, ranked search results, an embedding endpoint, health status and OpenAPI documentation. ZenML run metadata links operators to the deployed interface and supporting endpoints. A smoke profile is included for a smaller end-to-end test before running the full benchmark.
The example targets OpenShift Container Platform 4.20 or later and OpenShift AI 3.4 or later; its authors say it was validated with OpenShift AI 3.4.3 and ZenML 0.96.2. It requires the OpenShift AI KServe and MLflow components, the integrated image registry and persistent storage, but no GPU. The bootstrap needs cluster-administrator privileges or an equivalent permission set.
A pattern, not a production search platform
The repository is unusually explicit about its production-readiness limits. It describes the deployment as a proof of concept for isolated development and demonstrations, not production use. The sample loads a static exact FAISS index in the application process and does not provide continuous ingestion, metadata filtering or access control.
Its supporting services are also deliberately small-scale: ZenML, MySQL, MinIO, MLflow and model serving are not configured for high availability; external secret management, network policies, autoscaling, monitoring and backup procedures are outside the build.
Those caveats matter because the useful part of the project is the engineering pattern, not the sample topology. It shows how an OpenShift AI team can connect an evidence-based retrieval benchmark to a controlled deployment step, while keeping the evaluation run and its artifacts visible. Teams adopting it would still need to replace the demonstration storage, security and serving choices with production-grade services and controls.
The repository includes Helm charts, deployment and cleanup commands, validation targets, architecture diagrams and a small search UI. It is a concrete starting point for teams that want retrieval quality to be measured before an updated component reaches a RAG application or AI agent.
sources
comments · 0