Mosaic AI Vector Search and RAG for the Databricks GenAI Engineer Exam, Explained
Databricks GenAI Engineer · Concept Explained

Mosaic AI Vector Search and RAG for the Databricks GenAI Engineer exam, explained

Mosaic AI Vector Search is the managed vector database on Databricks: a serving endpoint hosts a Unity Catalog-governed index built from a source Delta table and an embedding model, and a RAG application queries it for the chunks that ground an answer. The Databricks Generative AI Engineer Associate exam gives it three objectives, plus every retrieval objective around them.

Last updated October 2026.

The four components

The exam objective reads "explain the key concepts and components of Mosaic AI Vector Search", and there are four. The endpoint is the serving compute; one endpoint can host several indexes. The index is the queryable object it hosts, living at catalog.schema.index and secured like any other Unity Catalog asset. The source Delta table holds the chunked text and its metadata, and the embedding model turns that text into the vectors the index stores and compares.

One governance detail recurs in questions: Unity Catalog grants and endpoint ACLs decide who can reach an index, but row-level security is not enforced inside a query. Per-user scoping is done with the filters argument at query time, set by the application from the user's context. Chapter 32 covers the parts and that caveat.

Delta Sync versus Direct Vector Access

The first design decision is who keeps the index current. A Delta Sync Index points at a source Delta table and Databricks runs the pipeline that keeps the vectors in step as rows change, either continuously (near real time, higher cost) or triggered on demand or a schedule. It needs Change Data Feed enabled on the source table (delta.enableChangeDataFeed = true), and a missing CDF is the classic reason a new sync fails. A Direct Vector Access Index has no source table and no automatic sync: you insert, update and delete vectors through the API yourself.

 Delta Sync, Databricks-computedDelta Sync, self-managedDirect Vector Access
SourceA Delta table with an embedding_source_columnA Delta table with a precomputed embedding_vector_columnNone; vectors are written through the API
Who computes vectorsVector Search, through the embedding model endpoint you nameYour own upstream jobYou
Who keeps it currentDatabricks, continuous or triggered syncDatabricks, continuous or triggered syncYou, with explicit upserts and deletes
Reach for it whenA governed corpus that should track its table with no code to runEmbedding already happens upstream in your pipelineStreaming writes, or you own the whole embedding pipeline

Index types and sync modes as taught in Chapter 32 of the Certified course, from the Databricks Vector Search documentation.

Standard versus storage-optimized endpoints

The second decision is the endpoint type, and the objective spells out the sizing signals: number of embeddings, update frequency, latency budget and cost. A scenario always gives enough of those to decide.

Standard

Latency in the tens of milliseconds, high queries per second, continuous or triggered sync, comfortable into the hundreds of millions of vectors. The fit for a live search bar or a chat assistant.

Storage optimized

Billions of vectors at a much lower cost per vector, answers in the hundreds of milliseconds, modest QPS, triggered sync only. The fit for a huge, cost-sensitive corpus with a tolerant latency budget.

The official guide's own sizing sample (about 80 queries per second over 100 million items, latency critical) lands on a standard endpoint with hybrid search and reranking switched off, because both add latency. Chapter 34 works that trade-off in both directions.

Where Vector Search sits in a RAG application

  1. Chunking. Documents are cut into the unit that gets embedded and retrieved. Chunks start every (size minus overlap) tokens, so raising the size and cutting the overlap both reduce the embedding count. Structure-aware, hierarchical and parent-document strategies take over when one fixed size cannot both match precisely and answer fully (Chapter 11, Chapter 12).
  2. Embedding. Chunks are written to a Delta table in Unity Catalog and embedded. The query must be embedded by the same model that built the index, or nearest-neighbour distances are noise.
  3. Retrieval. similarity_search takes query_text, num_results, columns and filters; columns decides which fields come back, filters decides who sees what. Plain ANN is the default and the fastest, and query_type="HYBRID" adds keyword matching when exact terms matter (Chapter 33).
  4. Re-ranking. A cross-encoder reads the query and each candidate together and re-scores a short list the first pass produced. It fixes the order at the top at the cost of latency, which is why latency-critical workloads run with it off (Chapter 15).
  5. Generation. The top chunks are stitched into the prompt and the LLM answers from them.

To serve that pipeline, MLflow packages six elements: the model flavor, the embedding model, the retriever, the declared dependencies, an input example and the signature inferred from it. Leave the index out of resources= and the deployed endpoint gets no credential to reach it (Chapter 28).

The retrieval metrics the exam uses

The course's stance, and the guide's, is measure before you tune. Precision@k asks what fraction of the k retrieved chunks are relevant; recall@k asks what fraction of the relevant chunks that exist made it into the top k. MRR scores where the first relevant hit landed, and NDCG credits every relevant result discounted by depth. At scale, mlflow.genai.evaluate() runs three retrieval judges over traces: RetrievalRelevance and RetrievalGroundedness need no ground truth, while RetrievalSufficiency needs expected answers (Chapter 14).

How the exam tests it

Section 4 (Assembling and Deploying Applications, 15 objectives) carries the three Vector Search objectives directly: explain its key concepts and components, create and query an index, and configure it from embedding count, update frequency, latency and cost. The same section asks for the basic elements of a RAG application and lists updating a Vector Search index among its CI/CD practices (Chapter 41).

Section 2 (Data Preparation, 8 objectives) tests the data side: chunking for a document structure, advanced chunking strategies, retrieval metrics and the role of re-ranking. Section 3 adds choosing an embedding model's context length and selecting a chunking strategy from retrieval evaluation (Chapter 22, Chapter 24). Across those sections, retrieval is the single largest theme on the paper.

FAQ

What is Mosaic AI Vector Search?

Mosaic AI Vector Search is the managed vector database built into Databricks. A serving endpoint hosts one or more indexes, each built from a source Delta table and an embedding model and governed as a Unity Catalog object, and applications query an index with similarity_search to retrieve the chunks that ground a RAG answer.

What is the difference between a Delta Sync index and a Direct Vector Access index?

A Delta Sync index points at a source Delta table and Databricks keeps it current automatically, with continuous or triggered sync, as long as Change Data Feed is enabled on the table. A Direct Vector Access index has no source table and no automatic sync, so you insert, update and delete vectors yourself through the API.

When should I choose a storage-optimized Vector Search endpoint over a standard one?

Choose storage optimized when the corpus runs to billions of vectors, cost matters more than speed, the refresh is batch rather than real time, and a latency of hundreds of milliseconds is acceptable. Choose standard for latency-critical, high-QPS workloads up to hundreds of millions of vectors, and for continuous sync, which storage optimized does not support.

Does the Databricks GenAI Engineer exam test retrieval metrics?

Yes, one Data Preparation objective is to use tools and metrics to evaluate retrieval performance. Know precision@k and recall@k, the ranking metrics MRR and NDCG, and the three MLflow retrieval judges, where RetrievalRelevance and RetrievalGroundedness need no ground truth and RetrievalSufficiency does.

How does re-ranking fit with Mosaic AI Vector Search?

Re-ranking is a query-time option that re-scores the candidates the first-pass vector search returned, reading the query and each chunk together so the best evidence rises to the top. It improves ordering at the cost of latency, so the guide's own sizing scenario turns it off for a latency-critical workload.

Study the Vector Search chapters in order

Chapters 32 to 34 of Certified's Generative AI Engineer Associate course cover the components, creating and querying an index, and the standard versus storage-optimized decision, each with quiz cards built from exam-style scenarios. Units 1 and 2 are free to start.

Open Chapter 32 →