Microsoft Azure AI Cloud Developer Associate (AI-200)

Use this AI-200 practice test to prepare for the Microsoft Azure AI Cloud Developer Associate exam. Original questions are grounded in verified Microsoft Learn content, with explanations and source links for focused revision.

Q001 - Question

A developer changed the Dockerfile for an inference API. The developer's workstation doesn't have Docker installed. You need to validate the Dockerfile change with a one-time build from the local directory before the change is committed to source control. You don't want to set up persistent automation. What should you use?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

D is correct.

Explanation: A quick task with az acr build uploads the local build context, builds the image in Azure, and pushes it to the registry. Quick tasks are intended for one-time builds, such as validating Dockerfile changes before you commit them, and don't require a local Docker installation.

A is incorrect: A source code trigger starts builds when code is committed or pull requests are created. The change isn't committed yet, and this option creates persistent automation.

B is incorrect: A scheduled trigger runs builds on a defined schedule. It doesn't provide an immediate build of uncommitted local files.

C is incorrect: az acr run with a /dev/null context runs a command in an existing image. It doesn't build a new image from your source files.

Learn more in Microsoft Learn

Q002 - Question

Your team plans to deploy three web apps to Azure App Service. Each app pulls container images from the same private Azure Container Registry (ACR). The platform team must configure registry permissions before the web apps are created. The team must not store registry credentials in App Service configuration. Which authentication approach should you use?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

C is correct.

Explanation: A user-assigned managed identity exists independently of the web app. You can create it, grant it the AcrPull role, and then assign it to one or more web apps. Managed identity authenticates to ACR with an Azure identity instead of stored credentials.

A is incorrect: A system-assigned managed identity is tied to the web app lifecycle. Azure creates it only when you enable it on an existing web app, so you can't grant permissions before the apps exist.

B is incorrect: Admin credentials store a username and password in your App Service configuration and require manual rotation if compromised.

D is incorrect: The Other container registries option uses a username and password for private images, so it also stores credentials in App Service configuration.

Learn more in Microsoft Learn

Q003 - Question

A containerized web app in Azure App Service has an API_KEY app setting that uses a Key Vault reference without a version specifier. The web app has a managed identity with access to read secrets. The security team rotates the secret and requires the app to use the new value now, not within the normal refresh window. What should you do?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

D is correct.

Explanation: A Key Vault reference without a version resolves to the latest secret version. App Service refreshes resolved values within 24 hours. Any configuration change that triggers an app restart forces an immediate refetch of referenced secrets.

A is incorrect: Storing the key in the image removes the Key Vault reference and requires a rebuild for each rotation, which app settings are designed to avoid.

B is incorrect: Slot settings remain with a deployment slot during swap operations. They don't control when App Service refetches a Key Vault secret.

C is incorrect: The continuous deployment webhook is called when new images are pushed to the registry. A secret rotation in Key Vault doesn't push an image.

Learn more in Microsoft Learn

Q004 - Question

A sidecar-enabled app pulls a model-server image from Azure Container Registry by using a user-assigned managed identity. The identity is assigned to the web app, has the applicable pull role, and the site container definition uses authType UserAssigned with the correct client ID. The container log shows an UNAUTHORIZED error that includes token validation failed. What should you check next?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

B is correct.

Explanation: Managed-identity image pulls require the registry to accept Azure Resource Manager audience tokens. An UNAUTHORIZED error with token validation failed can indicate that the registry doesn't accept these tokens, so verify the setting with az acr config authentication-as-arm show.

A is incorrect: An image pull failure occurs before the process starts. Don't use scaling to hide a configuration or authorization problem.

C is incorrect: DOCKER_REGISTRY_SERVER_* settings don't configure sidecar-enabled containers, and the solution must avoid stored registry passwords.

D is incorrect: The target port affects local communication after the process starts. It doesn't affect registry token validation during the image pull.

Learn more in Microsoft Learn

Q005 - Question

Your application image specifies a private custom PyTorch base image in its Dockerfile FROM statement. The application image has an ACR task with a base image trigger. The base image is stored in a different registry than the application image. When the base image is updated, the application image doesn't rebuild. What should you do to enable automatic trigger detection?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

A is correct.

Explanation: For private base images, store the base image and the application image in the same registry. This configuration is required for automatic trigger detection. After you move the base image, ACR can detect updates and rebuild the images that reference it.

B is incorrect: A source code trigger starts builds on commits or pull requests. It doesn't detect updates to a base image.

C is incorrect: The {{.Run.ID}} variable creates a unique tag for each build. It doesn't affect how ACR detects base image updates.

D is incorrect: A .dockerignore file reduces the context upload and can speed up builds. It doesn't enable base image update detection.

Learn more in Microsoft Learn

Q006 - Question

Your registry publishes inference-api with semantic version tags and stable tags. The current tags include inference-api:1, inference-api:1.1, inference-api:1.1.0, and inference-api:2.0.0. A consuming team wants to automatically receive backward-compatible features and bug fixes for version 1. The team must not receive breaking API changes. Which tag should the team reference?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

C is correct.

Explanation: The stable tag inference-api:1 points to the latest 1.x.x image. Consumers receive MINOR updates, which add backward-compatible features, and PATCH updates, which fix bugs. They don't receive MAJOR updates, which contain breaking changes.

A is incorrect: The latest tag moves whenever someone pushes a new image with that tag. It can include a new major version with breaking changes.

B is incorrect: inference-api:1.1.0 is a unique tag for one specific patch version. The team doesn't receive later features or fixes unless they update the reference.

D is incorrect: inference-api:2.0.0 is a new major version that contains a breaking API change.

Learn more in Microsoft Learn

Q007 - Question

You deploy a WebSocket server to Azure Container Apps. Clients keep connections open for long periods. The app must add replicas as the number of open connections grows. It must also scale to zero replicas when all connections close. Which scale rule should you configure?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

C is correct.

Explanation: TCP scaling adjusts the replica count based on concurrent TCP connections. Use it for WebSocket servers, gRPC services, and other workloads with long-lived connections. TCP scaling supports scale-to-zero. When all connections close, the app can scale to zero after the cool-down period.

A is incorrect: HTTP scaling counts requests received over a 15-second window. It fits short-lived request-response traffic, not persistent connections.

B is incorrect: CPU scale rules can't scale an app to zero. At least one replica must run so that utilization can be measured.

D is incorrect: Memory scale rules also require at least one running replica, so they can't meet the scale-to-zero requirement.

Learn more in Microsoft Learn

Q008 - Question

Your AI application has 400 registered users. During the busiest interval, about 120 users are online, and about 15 of them run code analysis at the same time. Each active conversation uses its own code interpreter session. You must limit cost and protect downstream resources. How should you configure the session pool's maximum concurrent sessions?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

A is correct.

Explanation: Base the maximum concurrent sessions value on expected concurrent active requests, not on total or online users. The limit also sets a boundary for cost and downstream resource use. Load testing shows how requests behave when the pool reaches that boundary.

B is incorrect: Total registered users overstate concurrency and remove the cost and resource protection that the limit provides.

C is incorrect: Online users aren't all running code. A value of 120 overestimates demand when only about 15 conversations execute code at once.

D is incorrect: Cooldown controls how long an idle session keeps its state. It doesn't limit how many sessions run at the same time, and a longer cooldown holds capacity longer.

Learn more in Microsoft Learn

Q009 - Question

After an update, upstream systems report that your AI API has intermittent delays and missing responses. You already know the name of the revision that serves traffic. You suspect that the revision scales to zero or has a crash loop. You need to verify the running instances of that specific revision. Which command should you run?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

B is correct.

Explanation: Replicas are the running instances of a revision. Use az containerapp replica list with --revision to check whether the specified revision runs and scales as expected. This check helps you detect scale-to-zero or crash loops during an incident.

A is incorrect: The registry list shows registries configured on the app. Use it to debug image pull failures, not to check running instances.

C is incorrect: The env show command shows environment details. Use it to confirm region and configuration, not replica state.

D is incorrect: The revision list shows active and inactive revisions. It confirms which version is active, but it doesn't show the running instances of a revision.

Learn more in Microsoft Learn

Q010 - Question

An upstream dependency for your AI background worker in Azure Container Apps has an outage. The failure isn't limited to one revision. While you mitigate the incident, you must ensure that the container app doesn't start any replicas. What should you do?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

A is correct.

Explanation: Stopping is an explicit app-level action. Use it when you need to ensure that the app doesn't start replicas while you mitigate an incident.

B is incorrect: Scale to zero depends on your scaling configuration, so it doesn't ensure that replicas stay stopped during the incident.

C is incorrect: Restart forces replicas to restart and can clear transient failure states, but replicas run again after the restart.

D is incorrect: Deactivation removes one revision from traffic. The problem affects the whole app, so other revisions can still run.

Learn more in Microsoft Learn

Q011 - Question

Your team stores a container app definition in a YAML file in source control and reviews all configuration changes like code. A developer needs to move the ai-api app to the image myregistry.azurecr.io/ai-api:v2. The change must stay consistent with the reviewed configuration. What should the developer do?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

A is correct.

Explanation: When you use --yaml, the YAML file is the source of truth. Change the image in the YAML file, review the change, and apply it with az containerapp update --yaml. The update creates a new revision that you can validate.

B is incorrect: When you use --yaml, other flags are ignored. The --image flag doesn't override the YAML file.

C is incorrect: Use az containerapp up for prototypes or a first deployment. It doesn't keep the change in the reviewed YAML file.

D is incorrect: Recreating the app without the YAML file bypasses code review and can cause configuration drift across environments.

Learn more in Microsoft Learn

Q012 - Question

Many container apps in your AI solution pull production images from Azure Container Registry. Rotating registry credentials across these apps takes significant effort. You need a predictable authentication method that reduces the number of secrets your team must rotate and follows least privilege. What should you do?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

D is correct.

Explanation: For Azure Container Registry, managed identity avoids long-lived credentials in deployment scripts. Grant the identity only the AcrPull role, and then configure the registry with --identity. Assign the identity and grant pull permissions before you use this pattern.

A is incorrect: Username and password works broadly, but it increases secret management overhead across many apps.

B is incorrect: Azure CLI can sometimes infer credentials, but this behavior isn't predictable. Use an explicit authentication approach in production.

C is incorrect: Least privilege requires that you assign only AcrPull to identities that pull images.

Learn more in Microsoft Learn

Q013 - Question

You deploy an AI inference API to Azure Kubernetes Service (AKS). The Pods stay in the Pending status and never transition to Running. You need to identify the cause and resolve the issue. What should you do?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

A is correct.

Explanation: Pending Pods usually indicate that Kubernetes cannot schedule the Pods due to insufficient resources. You should run kubectl describe pod to check for insufficient memory or CPU events, and then review resource requests or scale the cluster.

B is incorrect: Checking application logs is the solution for CrashLoopBackOff, which occurs when the application exits or crashes on startup.

C is incorrect: Checking for empty endpoints is the solution when a Service has no endpoints because the selector does not match the Pod labels.

D is incorrect: Verifying the image path and registry is the solution for ImagePullBackOff, which occurs when Kubernetes cannot pull the container image.

Learn more in Microsoft Learn

Q014 - Question

Your team runs a model inference API and a background data enrichment worker on Azure Kubernetes Service (AKS). When slowdowns occur, the team wants a quick visual assessment first. Then the team wants direct access to cluster resources, logs, and events for in-depth analysis. Which approach should you use?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

C is correct.

Explanation: The evidence states that the Azure portal offers visual tools like the Workloads blade, Live Logs, and the Diagnose and solve problems feature for quick inspection. For detailed investigation, kubectl commands give you direct access to cluster resources, logs, and events. Many developers use both approaches together.

A is incorrect: The Azure portal provides Live Logs to view container output, so it is not limited to non-log inspection.

B is incorrect: Diagnose and solve problems provides guided troubleshooting, but you still use kubectl for detailed investigation of resources, logs, and events.

D is incorrect: The goal of monitoring is to spot issues before they affect users. Redeploying does not identify whether the cause is application code, Kubernetes configuration, or the cluster.

Learn more in Microsoft Learn

Q015 - Question

A team runs a containerized Python FastAPI inference API on AKS as a single Pod that they created directly. They want Kubernetes to maintain a set number of copies of the API and restart any copy that crashes. What should they configure?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

B is correct.

Explanation: A Deployment defines how many replicas of your application run and specifies the container image and configuration. If a Pod crashes, the Deployment restarts it. You typically don't run single Pods directly because Deployments manage Pods for you.

A is incorrect: A ClusterIP Service provides a stable internal endpoint for Pods. It doesn't control how many Pods run or restart crashed Pods.

C is incorrect: kubectl logs shows log output for troubleshooting. It doesn't create or restart Pods.

D is incorrect: An ExternalName Service maps a Service to an external DNS name. It doesn't manage Pod replicas.

Learn more in Microsoft Learn

Q016 - Question

You deploy an AI inference API to Azure Kubernetes Service (AKS). The API needs environment-specific endpoint settings and API keys for upstream services. It also needs a location for user uploads that keeps data when Pods restart. You must not hardcode values into the container image. Which approach should you use?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

C is correct.

Explanation: Use ConfigMaps to separate nonsensitive configuration, such as endpoints, from code. Use Secrets to protect sensitive values, such as API keys. Use a PVC for durable storage that survives Pod restarts.

A is incorrect: API keys are sensitive values and belong in a Secret. Also, data in the container filesystem is lost when a Pod restarts.

B is incorrect: This option reverses the roles. API keys are sensitive and belong in a Secret, and endpoint settings are nonsensitive and belong in a ConfigMap.

D is incorrect: A PVC provides durable storage for data. It does not externalize settings or protect sensitive values the way ConfigMaps and Secrets do.

Learn more in Microsoft Learn

Q017 - Question

A production AI workload on AKS has strict compliance requirements. Secrets must remain exclusively in a dedicated vault. You need fine-grained audit trails for each secret access. Rotated secrets must reach the containers without Pod restarts. What should you use?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

A is correct.

Explanation: The Secrets Store CSI Driver mounts secrets from Key Vault into Pod filesystems as files. Secrets remain in Key Vault. The driver polls for changes and updates mounted secrets without Pod restarts. Choose direct Key Vault integration when you need fine-grained audit trails for each secret access.

B is incorrect: The App Configuration Kubernetes Provider generates native Kubernetes Secret resources with the resolved values. Use this option when you want to manage secret mappings centrally, not when secrets must remain exclusively in the vault.

C is incorrect: Kubernetes Secrets store the values in the cluster. Environment variable values from secretKeyRef are set at Pod start, so you must trigger a Deployment rollout after you update them.

D is incorrect: Never commit Secret manifests with literal values to source control. This option also does not keep secrets in a dedicated vault.

Learn more in Microsoft Learn

Q018 - Question

You run kubectl apply for an inference API Deployment that references myregistry.azurecr.io/inference-api:v1.0. The Pods show the status ImagePullBackOff. What should you do to diagnose and resolve the issue?

Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice

C is correct.

Explanation: ImagePullBackOff means Kubernetes can't pull your container image from the registry. Run kubectl describe pod and review the events section for the pull error. Then verify the image path, tag, and registry in your manifest, and confirm the image was built and pushed with az acr repository list.

A is incorrect: Resource requests affect scheduling. Requests that are too large cause Pending Pods, not image pull failures.

B is incorrect: Adding nodes with az aks scale resolves exhausted cluster capacity for Pending Pods. More nodes can't pull an image that doesn't exist or has the wrong path.

D is incorrect: A selector mismatch causes a Service to have no endpoints. It doesn't prevent the image from being pulled.

Learn more in Microsoft Learn

Q019 - Question

You build an AI-powered recommendation engine on Azure Cosmos DB for NoSQL. You use the default indexing policy. A new feature requires queries that sort products by more than one property with ORDER BY. What should you do to support this query pattern?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

C is correct.

Explanation: The default indexing policy automatically indexes the properties in your items, which reduces upfront index setup. As query patterns change, you might still need to customize the indexing policy. For example, some ORDER BY patterns require composite indexes.

A is incorrect: The partition key controls how data is distributed across physical partitions. It does not define index paths. You also cannot change a container's partition key without creating a new container and migrating data.

B is incorrect: Autoscale changes how RU/s scale with demand. It does not add the indexes that some ORDER BY patterns require.

D is incorrect: Shared throughput at the database level changes how capacity is allocated across containers. It does not change the indexing policy of a container.

Learn more in Microsoft Learn

Q020 - Question

Support agents use a semantic search application built on Azure Cosmos DB for NoSQL. Agents report that searches such as "error code 0x80070005 access denied" return conceptually related articles but often miss articles that contain the exact error code. You need results that reflect both exact term matches and semantic similarity in a single ranked list. What should you configure?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

D is correct.

Explanation: Hybrid search combines vector similarity with full-text scoring by using the RRF function. Configure a full-text policy and a full-text index on the text property, then rank with VectorDistance and FullTextScore so that the error code is matched exactly and "access denied" is matched semantically.

A is incorrect: A similarity threshold filters out low-quality semantic matches but doesn't add exact keyword matching for terms such as error codes.

B is incorrect: Multi-vector search combines scores from multiple embedding spaces. Both scores are still semantic, so exact term matching isn't added.

C is incorrect: float16 reduces storage, and a larger TOP N increases RU consumption. Neither change adds keyword matching.

Learn more in Microsoft Learn

Q021 - Question

A data ingestion pipeline writes large volumes of document chunks and embeddings to Azure Cosmos DB for NoSQL in real time. A reporting job queries the container once per day. The indexing policy includes /* and defines 12 composite indexes. Write RU consumption exceeds the budget. You need to reduce overall RU cost for this workload. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

B is correct.

Explanation: This workload is write-heavy. Every index increases write latency and RU consumption because indexes are updated synchronously during writes. Minimize indexed properties and composite indexes. When queries are infrequent, an occasional higher query cost can be acceptable compared to constantly elevated write costs.

A is incorrect: Each composite index adds write overhead. Creating composite indexes for edge cases increases write costs without proportional read benefits.

C is incorrect: Write RU costs don't vary by consistency level at the operation level. Strong consistency doubles read RU cost and increases write latency in multi-region accounts.

D is incorrect: Embedding arrays in range indexes increase storage and write overhead without query benefits. Each array element is indexed separately, and vector searches use vector indexes.

Learn more in Microsoft Learn

Q022 - Question

Your team has an existing Azure Cosmos DB for NoSQL account. An existing container stores support documents but was created without a vector policy. You need to store embeddings and run vector similarity queries on the support documents. Which two actions should you take? Each correct answer presents part of the solution.

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Multi-select

B and D are correct.

Explanation: You must enable vector search on the Azure Cosmos DB account before you use it. Vector policies must be set when the container is created, and vector search is only supported on new containers, so create a new container with both policies.

A is incorrect: Vector policies can't be modified after the container exists.

C is incorrect: Adding the embedding path to a range index doesn't configure vector search. Embedding arrays don't benefit from range indexes and increase write costs.

Learn more in Microsoft Learn

Q023 - Question

You are designing an Azure Cosmos DB for NoSQL container for a knowledge base. Each document stores a 3,072-dimension embedding from the text-embedding-3-large model. The container will grow to millions of vectors, and similarity queries must return results quickly with high accuracy. Which vector index type should you configure on the embedding path?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

C is correct.

Explanation: Use diskANN because it supports up to 4,096 dimensions and provides the best performance for large datasets with millions of vectors while maintaining high accuracy.

A is incorrect: The flat index performs exact brute-force search and limits vectors to 505 dimensions, so it can't index 3,072-dimension embeddings.

B is incorrect: quantizedFlat supports up to 4,096 dimensions but still performs brute-force search and is recommended for datasets up to approximately 50,000 vectors per physical partition.

D is incorrect: Embedding arrays don't benefit from standard range indexes. Including them wastes storage and increases write costs, and queries without a vector index perform a full scan.

Learn more in Microsoft Learn

Q024 - Question

A production RAG application queries an Azure Cosmos DB for NoSQL container that uses a DiskANN vector index. Before release, your team wants to validate how closely the indexed results match exact results for a set of test queries. Production queries must keep low latency and RU consumption. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

A is correct.

Explanation: Indexed search trades a small amount of accuracy for better performance. Use brute-force search for scenarios that require exact accuracy, such as evaluation and testing, and keep indexed search for production queries.

B is incorrect: Brute-force search compares the query vector against every document, consumes significantly more RUs, has higher latency, and scales poorly as data grows.

C is incorrect: Removing TOP N doesn't force exact search. Without TOP N, the query attempts to return all documents, which consumes excessive RUs and causes high latency.

D is incorrect: The distance function determines how similarity is calculated, not whether search is exact. Vector policies also can't be modified after the container exists.

Learn more in Microsoft Learn

Q025 - Question

A nightly Python job uses psycopg to load about 50,000 archived agent messages into the messages table in Azure Database for PostgreSQL. The job currently runs one INSERT statement per row in a loop and takes too long. You need the highest load throughput. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

B is correct.

Explanation: For larger datasets of 10,000 rows or more, the COPY command provides the highest performance. It is often two to 10 times faster than individual inserts. In psycopg, use cur.copy() and write each record with copy.write_row().

A is incorrect: Use executemany for inserting hundreds to a few thousand rows. For 50,000 rows, COPY provides higher throughput.

C is incorrect: Never use string formatting or concatenation to build queries with data values. Use parameterized queries to prevent SQL injection.

D is incorrect: Creating new database connections is expensive because each requires network handshakes, authentication, and server-side resource allocation. A new connection per row makes the load slower.

Learn more in Microsoft Learn

Q026 - Question

You design a products table for filtered vector search in Azure Database for PostgreSQL. Most queries filter by category_id and a price range before they order by vector similarity. Products also have vendor-specific attributes that differ between products and are rarely filtered. Which two design choices should you make? Each correct answer presents part of the solution. Select two.

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Multi-select

B and D are correct.

Explanation: Structured columns with a B-tree index on (category_id, price) let PostgreSQL narrow the candidate rows before it calculates vector distances. A JSONB column gives you schema flexibility for dynamic attributes that you rarely filter on. This combination optimizes the common case and keeps the design flexible.

A is incorrect: GIN indexes support containment queries but not range queries. A price range filter on JSONB requires an expression index or a sequential scan.

C is incorrect: Consider partitioning when tables exceed tens of millions of rows and queries filter by the partition key. Here, the queries filter mainly by category, so partitioning by price adds complexity and doesn't match the query pattern.

Learn more in Microsoft Learn

Q027 - Question

You store 1536-dimensional dense embeddings from text-embedding-ada-002 in a vector(1536) column in Azure Database for PostgreSQL. Storage costs are increasing as the document table grows. You must reduce embedding storage and keep acceptable similarity search quality. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

C is correct.

Explanation: The halfvec type stores elements as 16-bit floating-point numbers and uses half the storage of the vector type. Use halfvec only after you benchmark and verify that half precision provides acceptable search quality for your use case.

A is incorrect: The sparsevec type is for models that produce sparse embeddings where most elements are zero. Dense embeddings from text-embedding-ada-002 are not sparse.

B is incorrect: The vector(n) dimension must match the output dimension of the embedding model. A 768-dimension column causes insertion errors for 1536-dimension embeddings.

D is incorrect: An index does not reduce the size of the vector column. An IVFFlat index adds about 1 to 1.5 times the vector column size in storage.

Learn more in Microsoft Learn

Q028 - Question

Attorneys use a vector-only search over legal documents in Azure Database for PostgreSQL. When they search for a specific case name such as "Smith v. Jones," documents that contain that exact name sometimes do not appear in the results. You must return exact term matches and keep semantic relevance. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

B is correct.

Explanation: Hybrid search combines vector similarity with keyword-based full-text search to capture both semantic and lexical relevance. Use it when users search for names, codes, or terms that must match exactly. A GIN index on the full-text search vector speeds up keyword matching.

A is incorrect: A stricter distance threshold increases precision, but it still ranks only by semantic distance. It can remove more results and does not add exact term matching.

C is incorrect: Averaging vectors is for finding documents similar to several related examples. It does not add lexical matching for an exact case name.

D is incorrect: L2 distance is still a vector distance metric. Changing the operator does not add keyword matching, and most text embedding models are optimized for cosine similarity.

Learn more in Microsoft Learn

Q029 - Question

Your legal search application stores 1536-dimensional embeddings in an embedding column. You plan to move to a new embedding model that outputs 3,072 dimensions. The application must keep serving consistent results during the migration, and you must be able to roll back if the new model underperforms. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

D is correct.

Explanation: A parallel column strategy lets you populate new embeddings alongside the existing ones, compare search quality, and switch over when you are ready. This approach adds temporary storage overhead, but it avoids downtime and lets you roll back if the new model underperforms.

A is incorrect: Overwriting embeddings in place returns inconsistent results while the migration is in progress, and you lose the old vectors needed for rollback.

B is incorrect: You can't mix embeddings from different models in the same column. Different models produce vectors with different dimensions and semantic relationships.

C is incorrect: Rebuilding the index does not solve the dimension change or the need to keep the old embeddings available during validation and rollback.

Learn more in Microsoft Learn

Q030 - Question

Your team built a proof of concept for an AI agent on an Azure Database for PostgreSQL server that uses the Burstable tier. You are preparing the solution for production. The production workload has steady, predictable resource requirements, and the agent makes many short-lived database calls to store individual messages. You must use the built-in PgBouncer connection pooler. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

C is correct.

Explanation: Built-in PgBouncer is available only on the General Purpose and Memory Optimized tiers. General Purpose fits production workloads with steady, predictable resource requirements. After you enable PgBouncer, connect on port 6432 instead of 5432. You can change compute tiers after deployment with a brief restart.

A is incorrect: The Burstable tier doesn't support the built-in PgBouncer feature, so you can't enable it on this tier.

B is incorrect: Changing the port doesn't enable pooling. PgBouncer isn't available on the Burstable tier, so port 6432 doesn't route through a pooler.

D is incorrect: Port 5432 is the standard PostgreSQL port for direct connections. To use PgBouncer, connect on port 6432. Memory Optimized also targets large in-memory workloads, which this scenario doesn't require.

Learn more in Microsoft Learn

Q031 - Question

A retailer generates image embeddings from a vision model for two million product photos. The embeddings aren't pre-normalized. The similarity search must return results in under 10 ms, and the team wants to keep Redis memory costs low. Which vector field configuration should you use?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

D is correct.

Explanation: Use FLOAT32 because it uses 4 bytes per dimension and provides enough precision for AI embeddings. Use L2 for image embeddings and spatial data. Use HNSW for datasets over 10,000 vectors when you need queries under 10 ms and can accept 95-99% accuracy.

A is incorrect: FLOAT64 uses 8 bytes per dimension, which doubles memory usage and slows distance calculations without a meaningful accuracy gain for embeddings.

B is incorrect: COSINE is the metric for text embeddings, and FLAT compares the query against every vector. FLAT query time grows linearly, so it doesn't meet the latency target at two million vectors.

C is incorrect: Use IP only with pre-normalized embeddings or specialized models that require it. These embeddings aren't pre-normalized.

Learn more in Microsoft Learn

Q032 - Question

Your team prepares a workstation for a lab task. The task creates an Azure Managed Redis resource by using the Azure CLI. It then runs a Python app that loads, stores, and searches vector data with metadata. Which preparation meets the documented prerequisites for this task?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

B is correct.

Explanation: To create the resource and run the app, you need an Azure subscription with permission to create an Azure Managed Redis instance with an enterprise SKU. You also need Visual Studio Code, Python 3.12 or greater, the latest version of the Azure CLI, and the Azure CLI redisenterprise extension.

A is incorrect: Python 3.10 doesn't meet the Python 3.12 or greater requirement, and this option omits the redisenterprise extension and the subscription permission.

C is incorrect: The Azure CLI alone isn't enough. You must also install the redisenterprise extension.

D is incorrect: Read-only access doesn't let you create the Azure Managed Redis instance. You need permission to create an instance with an enterprise SKU.

Learn more in Microsoft Learn

Q033 - Question

You develop a Python AI application that uses the redis-py library. The Azure Managed Redis instance uses the OSS clustering policy. You need to configure the client connection correctly for this clustering policy. What should you use?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

D is correct.

Explanation: The clustering policy chosen for an Azure Managed Redis instance impacts the connection method. If you are working with an instance using the OSS clustering policy, you need to use redis.cluster.RedisCluster instead of redis.Redis for your connection.

A is incorrect: Port 6380 is the default for Azure Cache for Redis, not Azure Managed Redis (which uses 10000). Changing the port does not configure the client for the OSS clustering policy.

B is incorrect: The decode_responses=False parameter configures the client to work with raw bytes instead of strings. It does not configure the client for the OSS clustering policy.

C is incorrect: A credential_provider configures Microsoft Entra ID authentication. It does not configure the client for the OSS clustering policy.

Learn more in Microsoft Learn

Q034 - Question

You design vector storage for a product catalog in Azure Managed Redis. Each product has nested specifications, several variant records, and two embeddings: one for the description and one for the main image. The application already exchanges product data in JSON format. Which storage approach should you use?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

A is correct.

Explanation: Choose Redis JSON when your data has nested structures, you need multiple vectors per item, or your application already uses JSON. Convert each embedding to a numeric array with tolist(), and define the index with IndexType.JSON, $. JSONPath field names, and the as_name parameter for query references.

B is incorrect: Redis Hash stores flat field-value pairs only. It's the better choice for simple records when you need maximum memory efficiency and query speed, but it doesn't support nested objects.

C is incorrect: Flattening nested data into one text field works against the documented guidance. Hash is intended for data models that won't need nested objects.

D is incorrect: JSON documents store vectors as numeric arrays, not binary bytes, and a JSON-based index must use IndexType.JSON.

Learn more in Microsoft Learn

Q035 - Question

A web application for an AI shopping assistant stores shopping carts, user preferences, and authentication tokens in browser cookies. Bandwidth usage and latency increase as session data grows. You need to reduce request overhead and keep fast access to session data by using Azure Managed Redis. What should you do?

Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice

A is correct.

Explanation: Storing too much information directly in cookies negatively impacts performance because each cookie is transmitted with every HTTP request and response. The typical solution stores only a session identifier in the cookie, then uses that key to retrieve full session data from Azure Managed Redis, which delivers sub-millisecond latency.

B is incorrect: Keeping the full session data in the cookie means every request still transmits the data, which does not reduce bandwidth usage or latency.

C is incorrect: Caching static headers and footers is a content cache pattern, not a session store pattern. Querying session data from a backend database is slower than retrieving it from Azure Managed Redis.

D is incorrect: Compressing the session data in the cookie still transmits the data with every request and response, and this approach does not use Azure Managed Redis as a session store.

Learn more in Microsoft Learn

Q036 - Question

Your image classification pipeline uses a Service Bus Standard tier namespace. Each request includes an image of about 2 MB. The business doesn't want to upgrade to the Premium tier only for message size. The security team also wants access to the image data controlled separately from access to the messaging layer. What should you do?

Domain: Connect to and consume Azure services (20–25%) Type: Single choice

A is correct.

Explanation: Use the claim-check pattern. The Standard tier supports messages up to 256 KB. With claim-check, the message carries only a reference, so it works within any tier's size limits. You can also scope Blob Storage access independently from Service Bus access.

B is incorrect: Batching reduces network round trips. The SDK keeps a batch within the maximum message size for your tier, so batching doesn't let a single 2-MB message fit.

C is incorrect: The content_type property only describes the encoding format. A base64-encoded 2-MB image still exceeds the 256-KB Standard tier limit.

D is incorrect: The time_to_live property controls how long a message stays in the queue before it expires. It has no effect on the maximum message size.

Learn more in Microsoft Learn

Q037 - Question

An orchestration processes a large claim archive with one flat fan-out/fan-in pattern. Measurements show that the fan-in step is slow because one orchestrator loads and aggregates every individual document result. You need to reduce the state that one orchestration handles. You also want an independent status for each part of the archive. What should you do?

Domain: Connect to and consume Azure services (20–25%) Type: Single choice

D is correct.

Explanation: One orchestrator performs the fan-in step on one worker at a time, so a flat fan-in can become a bottleneck. Sub-orchestrations let child orchestrations process separate partitions and return compact summaries. The parent aggregates those summaries instead of every document result. Each partition also gets an independent status.

A is incorrect: Sequential processing adds the latency of every item and doesn't reduce the results that one orchestrator must aggregate.

B is incorrect: Returning extracted text expands orchestration history, and Durable Functions must load that data during replay.

C is incorrect: Scheduling more work at once doesn't reduce the aggregation work in one orchestrator. A sudden activity burst can also exceed model rate limits or saturate a downstream data service.

Learn more in Microsoft Learn

Q038 - Question

You maintain a Durable Functions app that processes insurance claim documents. The orchestrator passes full extracted document text to activities. Activities return complete model responses to the orchestrator. Replay time and storage operations are increasing. The orchestration history also retains sensitive claim content. What should you change?

Domain: Connect to and consume Azure services (20–25%) Type: Single choice

B is correct.

Explanation: Durable Functions serializes orchestration inputs, activity inputs, and activity outputs into the durable store. Large payloads increase storage operations, replay time, memory use, and the amount of sensitive data retained in history. Store large content in a data service such as Azure Blob Storage. Pass only a reference and the fields that the orchestrator needs for its next decision.

A is incorrect: Orchestration inputs are also serialized into orchestration history, so the full text stays in the durable store.

C is incorrect: The orchestrator still receives and handles the complete model responses. Keep large results out of the orchestration, write them from the activity, and return a reference.

D is incorrect: Don't place secrets or access tokens in orchestration inputs or outputs. Use managed identities and a service such as Azure Key Vault for credentials.

Learn more in Microsoft Learn

Q039 - Question

A Service Bus-triggered function calls Azure AI Document Intelligence to extract text from each document. Monitoring shows extra latency on every invocation because the function creates a new DocumentIntelligenceClient and credential each time it runs. You need to reduce this per-invocation overhead. What should you do?

Domain: Connect to and consume Azure services (20–25%) Type: Single choice

B is correct.

Explanation: You should initialize SDK clients outside the function handler at the module level. The objects persist across invocations on the same instance, so you avoid repeated connection setup, configuration loading, and token acquisition on every function invocation.

A is incorrect: Azure AI Document Intelligence does not have a dedicated binding. You must create SDK clients directly in your function code for this service.

C is incorrect: The host.json file contains global runtime settings such as timeouts, logging, and trigger concurrency. It does not create SDK clients for your code.

D is incorrect: The maxConcurrentCalls setting controls how many messages each instance processes at the same time. It does not remove the cost of creating a client on every invocation.

Learn more in Microsoft Learn

Q040 - Question

A function app has a system-assigned managed identity. One function writes classification results by using a Cosmos DB output binding configured with CosmosDBConnection__accountEndpoint. The same function calls Azure OpenAI Service through an SDK client that uses DefaultAzureCredential. You need to remove all connection strings and API keys and apply least privilege. Which two role assignments should you grant to the managed identity? Each correct answer presents part of the solution.

Domain: Connect to and consume Azure services (20–25%) Type: Multiple choice

B and C are correct.

Explanation: An identity-based Cosmos DB connection uses the account endpoint, and the managed identity needs the Cosmos DB Built-in Data Contributor role to write documents. Azure OpenAI Service has no dedicated binding, so the SDK client authenticates with DefaultAzureCredential. In production, this credential uses the managed identity, which needs the Cognitive Services OpenAI User role.

A is incorrect: The Key Vault Secrets User role is required for Key Vault references. This design removes stored secrets, so the function does not read from Key Vault.

D is incorrect: The Azure Service Bus Data Sender role is for Service Bus output bindings. This function does not write to Service Bus, so the role adds access the function does not need.

Learn more in Microsoft Learn

Q041 - Question

Your team runs a monitoring tool that checks whether the secrets that a RAG pipeline requires exist in Azure Key Vault. The tool must read secret names and properties. It must not read secret values. Which built-in role should you assign to the tool's identity?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

B is correct.

Explanation: Key Vault Reader grants read access to vault metadata, such as secret names and properties. It doesn't reveal secret values or key material. Use it for monitoring and discovery tools that verify which secrets exist.

A is incorrect: Key Vault Secrets User grants read access to secret values. The tool must not read values, so this role grants more access than the tool requires.

C is incorrect: Key Vault Contributor is a control plane role. It manages the vault resource, but it doesn't grant data plane operations on secrets, keys, or certificates.

D is incorrect: Key Vault Secrets Officer grants full management permissions on secrets, including create, update, and delete. Assign it to operators or CI/CD pipelines that manage the secret lifecycle, not to a read-only monitoring tool.

Learn more in Microsoft Learn

Q042 - Question

You add error handling to a Python service that calls get_secret() on a SecretClient. You want the service to retry only failures that might resolve without operator intervention. Missing secrets and missing RBAC role assignments must fail immediately so that operators can diagnose them. Which exception type should trigger a retry?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

C is correct.

Explanation: ServiceRequestError indicates a network-level problem, such as a DNS resolution failure or a timeout. This type of failure might resolve on retry.

A is incorrect: ResourceNotFoundError means that the secret name doesn't exist in the vault. This is a configuration error that requires operator intervention, so a retry won't resolve it.

B is incorrect: HttpResponseError covers authentication and authorization failures, such as an identity that lacks the required RBAC role. Retrying doesn't fix a missing role assignment.

D is incorrect: Retrying every exception also retries configuration and authorization errors. These errors won't resolve on their own, and retries delay diagnosis.

Learn more in Microsoft Learn

Q043 - Question

Your AI pipeline needs three settings: Pipeline:BatchSize, OpenAI:DeploymentName, and Storage:AccountKey. The storage account key must rotate on a schedule, and you need an audit trail of each access. The app must retrieve all three settings through a single load() call. Which design should you use?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

B is correct.

Explanation: Values that grant access to a resource belong in Key Vault, which provides rotation, expiration policies, and per-object audit logging. Nonsensitive values such as batch sizes and deployment names belong in App Configuration. A Key Vault reference keeps App Configuration as the single entry point, so one load() call returns all three settings.

A is incorrect: App Configuration doesn't provide per-object audit logging, HSM-backed encryption, expiration policies, or automated rotation for the account key.

C is incorrect: Key Vault has stricter throttling limits and doesn't support labels, feature flags, or snapshots for nonsensitive settings. This design also doesn't use a single load() call.

D is incorrect: Independent copies in two services can drift apart, and the app might read a stale value. A Key Vault reference stores only a pointer, so there's one source of truth.

Learn more in Microsoft Learn

Q044 - Question

A Python app uses a managed identity and calls load() with both credential and keyvault_credential set to DefaultAzureCredential. The identity has the App Configuration Data Reader role on the store. Regular settings load, but the provider throws an error that names the Key Vault and the secret for OpenAI:ApiKey. What should you do to resolve the error?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

D is correct.

Explanation: To resolve Key Vault references, the app identity needs a role on each service: App Configuration Data Reader on the store and Key Vault Secrets User on each referenced vault. When Key Vault access is missing, the provider throws an error that identifies the vault and secret it couldn't access.

A is incorrect: The identity already reads the store. App Configuration Data Owner adds write access to the store but doesn't grant access to secrets in Key Vault.

B is incorrect: Storing the secret directly in App Configuration removes Key Vault protections such as audit logging, rotation, and HSM-backed encryption.

C is incorrect: Pinning a version doesn't fix a missing permission, and it stops the app from picking up rotated secrets without a configuration change.

Learn more in Microsoft Learn

Q045 - Question

Operators often change several related settings in Azure App Configuration at the same time, such as Pipeline:BatchSize and Pipeline:RetryCount. The Python app must use the new values without a restart, and operators must control when the changes take effect so that the app doesn't apply a partial update. What should you implement?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

A is correct.

Explanation: With the sentinel key pattern, operators change the settings first and then update the Sentinel key. When the provider detects the change to the watched key, it reloads the entire configuration, so the app receives all changes together on the next refresh cycle.

B is incorrect: The provider doesn't push values. You must call refresh() explicitly, and you need refresh_on to define the sentinel key to watch.

C is incorrect: A restart doesn't meet the requirement to pick up changes without restarting the app.

D is incorrect: Redeploying the app for each change removes the benefit of updating configuration without redeployment.

Learn more in Microsoft Learn

Q046 - Question

You publish an Azure dashboard that contains pinned log query tiles from the pipeline's Application Insights resource. A team member has read permission on the dashboard resource. When the team member opens the dashboard, the log query tiles display an access error. What should you do?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

D is correct.

Explanation: Dashboard access uses Azure role-based access control (RBAC). Team members need read permissions on both the dashboard resource and the underlying data sources. Without access to the Application Insights resource, tiles that query that resource display an access error.

A is incorrect: Publishing makes the dashboard available to other users, but it doesn't grant access to the underlying Application Insights data.

B is incorrect: A Markdown tile adds context to the dashboard. It doesn't change permissions on the data source.

C is incorrect: The render operator preserves the visualization when you pin results. It doesn't resolve a permissions error.

Learn more in Microsoft Learn

Q047 - Question

Four Python services in a RAG pipeline use the Azure Monitor OpenTelemetry Distro and send telemetry to the same Application Insights resource. On the Application Map, all four services appear as a single node. You need each service to appear as a separate node while keeping the services logically grouped. What should you configure?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

A is correct.

Explanation: Application Insights forms the cloud role name from service.namespace combined with service.name. A unique service.name per service creates a separate Application Map node for each service. A shared service.namespace groups the services together logically.

B is incorrect: service.instance.id distinguishes multiple instances of the same service. With the same service.name, the services still share one cloud role name.

C is incorrect: The connection string tells the exporter which Application Insights resource receives telemetry. It does not set the cloud role name, and the services need to send to the same resource.

D is incorrect: Span attributes describe a single operation and appear in customDimensions. The cloud role name comes from resource attributes, which apply to all telemetry from a service.

Learn more in Microsoft Learn

Q048 - Question

Your client requires two monitoring outcomes for the content moderation pipeline. First, the team must be notified when the 95th-percentile response time for any service exceeds three seconds. Second, the team must be notified about gradual performance degradation that might never exceed a fixed threshold. No one watches the monitoring views continuously. Which two solutions should you implement? Each correct answer presents part of the solution.

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Multiple choice

B and D are correct.

Explanation: A log search alert can use any KQL query, including percentile calculations, so it enforces the specific three-second p95 threshold. Smart detection learns the normal behavior of your application and detects performance that is worse than its historical norm, including gradual degradation that never exceeds a fixed number.

A is incorrect: An average response time can stay within acceptable limits even when a significant percentage of requests is slow. It doesn't enforce a p95 requirement.

C is incorrect: A dashboard tile shows latency, but it requires someone to be watching. It doesn't notify the team.

Learn more in Microsoft Learn

Q049 - Question

You query the dependencies table for failed calls from the classification service. Most failures target the model inference endpoint and have a resultCode of 429. What should you do to address these failures?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

B is correct.

Explanation: An HTTP 429 result code from a model inference endpoint indicates rate limiting. You address rate limiting by adjusting throughput or implementing retry policies.

A is incorrect: A server-side failure in the dependency is indicated by an HTTP 500 result code, not 429.

C is incorrect: Network issues or an overloaded service typically appear as a timeout with no result code. These failures have a result code of 429.

D is incorrect: The failed calls originate from the classification service and target the inference endpoint. The ingestion API isn't the source of these dependency failures.

Learn more in Microsoft Learn

Q050 - Question

You monitor a retrieval-augmented generation (RAG) pipeline that runs as four microservices. Your client requires 95th-percentile response times under three seconds. You need to detect when response times trend above this target and trigger an alert. Which observability pillar should you use for this requirement?

Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice

C is correct.

Explanation: Metrics provide aggregate numerical measurements over time, such as request counts, error rates, and response-time percentiles. Use metrics to detect trends and set alerting thresholds for service-level objectives, such as a 95th-percentile target.

A is incorrect: Distributed tracing shows the path and timing of individual requests. Use traces to find where a problem occurs after metrics show that something changed.

B is incorrect: Logs are timestamped records of discrete events within a service. Logs explain why a specific operation behaved a certain way, but they do not provide aggregate percentile measurements for alert thresholds.

D is incorrect: Context propagation carries trace and span IDs across service boundaries. It is a mechanism that connects spans into one trace, not an observability pillar for percentile alerting.

Learn more in Microsoft Learn

← Read the full AI-200 study guide