Mosaic AI Vector Search is the managed vector database on Databricks: a serving endpoint hosts a Unity Catalog-governed index built from a source Delta table and an embedding model, and a RAG application queries it for the chunks that ground an answer. The Databricks Generative AI Engineer Associate exam gives it three objectives, plus every retrieval objective around them.
Last updated October 2026.
The exam objective reads "explain the key concepts and components of Mosaic AI Vector Search", and there are four. The endpoint is the serving compute; one endpoint can host several indexes. The index is the queryable object it hosts, living at catalog.schema.index and secured like any other Unity Catalog asset. The source Delta table holds the chunked text and its metadata, and the embedding model turns that text into the vectors the index stores and compares.
One governance detail recurs in questions: Unity Catalog grants and endpoint ACLs decide who can reach an index, but row-level security is not enforced inside a query. Per-user scoping is done with the filters argument at query time, set by the application from the user's context. Chapter 32 covers the parts and that caveat.
The first design decision is who keeps the index current. A Delta Sync Index points at a source Delta table and Databricks runs the pipeline that keeps the vectors in step as rows change, either continuously (near real time, higher cost) or triggered on demand or a schedule. It needs Change Data Feed enabled on the source table (delta.enableChangeDataFeed = true), and a missing CDF is the classic reason a new sync fails. A Direct Vector Access Index has no source table and no automatic sync: you insert, update and delete vectors through the API yourself.
| Delta Sync, Databricks-computed | Delta Sync, self-managed | Direct Vector Access | |
|---|---|---|---|
| Source | A Delta table with an embedding_source_column | A Delta table with a precomputed embedding_vector_column | None; vectors are written through the API |
| Who computes vectors | Vector Search, through the embedding model endpoint you name | Your own upstream job | You |
| Who keeps it current | Databricks, continuous or triggered sync | Databricks, continuous or triggered sync | You, with explicit upserts and deletes |
| Reach for it when | A governed corpus that should track its table with no code to run | Embedding already happens upstream in your pipeline | Streaming writes, or you own the whole embedding pipeline |
Index types and sync modes as taught in Chapter 32 of the Certified course, from the Databricks Vector Search documentation.
The second decision is the endpoint type, and the objective spells out the sizing signals: number of embeddings, update frequency, latency budget and cost. A scenario always gives enough of those to decide.
Latency in the tens of milliseconds, high queries per second, continuous or triggered sync, comfortable into the hundreds of millions of vectors. The fit for a live search bar or a chat assistant.
Billions of vectors at a much lower cost per vector, answers in the hundreds of milliseconds, modest QPS, triggered sync only. The fit for a huge, cost-sensitive corpus with a tolerant latency budget.
The official guide's own sizing sample (about 80 queries per second over 100 million items, latency critical) lands on a standard endpoint with hybrid search and reranking switched off, because both add latency. Chapter 34 works that trade-off in both directions.
similarity_search takes query_text, num_results, columns and filters; columns decides which fields come back, filters decides who sees what. Plain ANN is the default and the fastest, and query_type="HYBRID" adds keyword matching when exact terms matter (Chapter 33).To serve that pipeline, MLflow packages six elements: the model flavor, the embedding model, the retriever, the declared dependencies, an input example and the signature inferred from it. Leave the index out of resources= and the deployed endpoint gets no credential to reach it (Chapter 28).
The course's stance, and the guide's, is measure before you tune. Precision@k asks what fraction of the k retrieved chunks are relevant; recall@k asks what fraction of the relevant chunks that exist made it into the top k. MRR scores where the first relevant hit landed, and NDCG credits every relevant result discounted by depth. At scale, mlflow.genai.evaluate() runs three retrieval judges over traces: RetrievalRelevance and RetrievalGroundedness need no ground truth, while RetrievalSufficiency needs expected answers (Chapter 14).
Section 4 (Assembling and Deploying Applications, 15 objectives) carries the three Vector Search objectives directly: explain its key concepts and components, create and query an index, and configure it from embedding count, update frequency, latency and cost. The same section asks for the basic elements of a RAG application and lists updating a Vector Search index among its CI/CD practices (Chapter 41).
Section 2 (Data Preparation, 8 objectives) tests the data side: chunking for a document structure, advanced chunking strategies, retrieval metrics and the role of re-ranking. Section 3 adds choosing an embedding model's context length and selecting a chunking strategy from retrieval evaluation (Chapter 22, Chapter 24). Across those sections, retrieval is the single largest theme on the paper.
Mosaic AI Vector Search is the managed vector database built into Databricks. A serving endpoint hosts one or more indexes, each built from a source Delta table and an embedding model and governed as a Unity Catalog object, and applications query an index with similarity_search to retrieve the chunks that ground a RAG answer.
A Delta Sync index points at a source Delta table and Databricks keeps it current automatically, with continuous or triggered sync, as long as Change Data Feed is enabled on the table. A Direct Vector Access index has no source table and no automatic sync, so you insert, update and delete vectors yourself through the API.
Choose storage optimized when the corpus runs to billions of vectors, cost matters more than speed, the refresh is batch rather than real time, and a latency of hundreds of milliseconds is acceptable. Choose standard for latency-critical, high-QPS workloads up to hundreds of millions of vectors, and for continuous sync, which storage optimized does not support.
Yes, one Data Preparation objective is to use tools and metrics to evaluate retrieval performance. Know precision@k and recall@k, the ranking metrics MRR and NDCG, and the three MLflow retrieval judges, where RetrievalRelevance and RetrievalGroundedness need no ground truth and RetrievalSufficiency does.
Re-ranking is a query-time option that re-scores the candidates the first-pass vector search returned, reading the query and each chunk together so the best evidence rises to the top. It improves ordering at the cost of latency, so the guide's own sizing scenario turns it off for a latency-critical workload.
Chapters 32 to 34 of Certified's Generative AI Engineer Associate course cover the components, creating and querying an index, and the standard versus storage-optimized decision, each with quiz cards built from exam-style scenarios. Units 1 and 2 are free to start.
Open Chapter 32 →