Q001 - Question
A developer changed the Dockerfile for an inference API. The developer's workstation doesn't have Docker installed. You need to validate the Dockerfile change with a one-time build from the local directory before the change is committed to source control. You don't want to set up persistent automation. What should you use?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Create an ACR task with az acr task create and a source code trigger on the main branch.
- B. Create an ACR task with az acr task create and a --schedule cron expression.
- C. Run az acr run with a /dev/null context to execute the image.
- D. Run a quick task with az acr build and the local directory as the build context.
D is correct.
Explanation: A quick task with az acr build uploads the local build context, builds the image in Azure, and pushes it to the registry. Quick tasks are intended for one-time builds, such as validating Dockerfile changes before you commit them, and don't require a local Docker installation.
A is incorrect: A source code trigger starts builds when code is committed or pull requests are created. The change isn't committed yet, and this option creates persistent automation.
B is incorrect: A scheduled trigger runs builds on a defined schedule. It doesn't provide an immediate build of uncommitted local files.
C is incorrect: az acr run with a /dev/null context runs a command in an existing image. It doesn't build a new image from your source files.
Q002 - Question
Your team plans to deploy three web apps to Azure App Service. Each app pulls container images from the same private Azure Container Registry (ACR). The platform team must configure registry permissions before the web apps are created. The team must not store registry credentials in App Service configuration. Which authentication approach should you use?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Enable a system-assigned managed identity on each web app and grant each identity the AcrPull role.
- B. Enable the admin user on the registry and use the admin credentials in each web app.
- C. Create a user-assigned managed identity, grant it the AcrPull role on the registry, and assign it to each web app.
- D. Select Other container registries and enter the ACR server URL, username, and password for each web app.
C is correct.
Explanation: A user-assigned managed identity exists independently of the web app. You can create it, grant it the AcrPull role, and then assign it to one or more web apps. Managed identity authenticates to ACR with an Azure identity instead of stored credentials.
A is incorrect: A system-assigned managed identity is tied to the web app lifecycle. Azure creates it only when you enable it on an existing web app, so you can't grant permissions before the apps exist.
B is incorrect: Admin credentials store a username and password in your App Service configuration and require manual rotation if compromised.
D is incorrect: The Other container registries option uses a username and password for private images, so it also stores credentials in App Service configuration.
Q003 - Question
A containerized web app in Azure App Service has an API_KEY app setting that uses a Key Vault reference without a version specifier. The web app has a managed identity with access to read secrets. The security team rotates the secret and requires the app to use the new value now, not within the normal refresh window. What should you do?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Rebuild the container image with the new API key and update the image tag.
- B. Mark API_KEY as a slot setting.
- C. Enable continuous deployment so the registry webhook refreshes the secret.
- D. Make a configuration change that triggers an app restart.
D is correct.
Explanation: A Key Vault reference without a version resolves to the latest secret version. App Service refreshes resolved values within 24 hours. Any configuration change that triggers an app restart forces an immediate refetch of referenced secrets.
A is incorrect: Storing the key in the image removes the Key Vault reference and requires a rebuild for each rotation, which app settings are designed to avoid.
B is incorrect: Slot settings remain with a deployment slot during swap operations. They don't control when App Service refetches a Key Vault secret.
C is incorrect: The continuous deployment webhook is called when new images are pushed to the registry. A secret rotation in Key Vault doesn't push an image.
Q004 - Question
A sidecar-enabled app pulls a model-server image from Azure Container Registry by using a user-assigned managed identity. The identity is assigned to the web app, has the applicable pull role, and the site container definition uses authType UserAssigned with the correct client ID. The container log shows an UNAUTHORIZED error that includes token validation failed. What should you check next?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Scale the App Service plan to a larger SKU so that the image pull has more resources.
- B. Run az acr config authentication-as-arm show to verify that the registry accepts Azure Resource Manager audience tokens.
- C. Enable the registry admin account and store its credentials in DOCKER_REGISTRY_SERVER_* app settings.
- D. Change the model-server target port so that it doesn't conflict with the main container.
B is correct.
Explanation: Managed-identity image pulls require the registry to accept Azure Resource Manager audience tokens. An UNAUTHORIZED error with token validation failed can indicate that the registry doesn't accept these tokens, so verify the setting with az acr config authentication-as-arm show.
A is incorrect: An image pull failure occurs before the process starts. Don't use scaling to hide a configuration or authorization problem.
C is incorrect: DOCKER_REGISTRY_SERVER_* settings don't configure sidecar-enabled containers, and the solution must avoid stored registry passwords.
D is incorrect: The target port affects local communication after the process starts. It doesn't affect registry token validation during the image pull.
Q005 - Question
Your application image specifies a private custom PyTorch base image in its Dockerfile FROM statement. The application image has an ACR task with a base image trigger. The base image is stored in a different registry than the application image. When the base image is updated, the application image doesn't rebuild. What should you do to enable automatic trigger detection?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Store the private base image and the application image in the same registry.
- B. Replace the base image trigger with a source code trigger on the main branch.
- C. Change the application image tag to use the {{.Run.ID}} run variable.
- D. Add a .dockerignore file to reduce the size of the build context.
A is correct.
Explanation: For private base images, store the base image and the application image in the same registry. This configuration is required for automatic trigger detection. After you move the base image, ACR can detect updates and rebuild the images that reference it.
B is incorrect: A source code trigger starts builds on commits or pull requests. It doesn't detect updates to a base image.
C is incorrect: The {{.Run.ID}} variable creates a unique tag for each build. It doesn't affect how ACR detects base image updates.
D is incorrect: A .dockerignore file reduces the context upload and can speed up builds. It doesn't enable base image update detection.
Q006 - Question
Your registry publishes inference-api with semantic version tags and stable tags. The current tags include inference-api:1, inference-api:1.1, inference-api:1.1.0, and inference-api:2.0.0. A consuming team wants to automatically receive backward-compatible features and bug fixes for version 1. The team must not receive breaking API changes. Which tag should the team reference?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. inference-api:latest
- B. inference-api:1.1.0
- C. inference-api:1
- D. inference-api:2.0.0
C is correct.
Explanation: The stable tag inference-api:1 points to the latest 1.x.x image. Consumers receive MINOR updates, which add backward-compatible features, and PATCH updates, which fix bugs. They don't receive MAJOR updates, which contain breaking changes.
A is incorrect: The latest tag moves whenever someone pushes a new image with that tag. It can include a new major version with breaking changes.
B is incorrect: inference-api:1.1.0 is a unique tag for one specific patch version. The team doesn't receive later features or fixes unless they update the reference.
D is incorrect: inference-api:2.0.0 is a new major version that contains a breaking API change.
Q007 - Question
You deploy a WebSocket server to Azure Container Apps. Clients keep connections open for long periods. The app must add replicas as the number of open connections grows. It must also scale to zero replicas when all connections close. Which scale rule should you configure?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. An HTTP scale rule that uses --scale-rule-http-concurrency
- B. A CPU scale rule with a Utilization value of 70 percent
- C. A TCP scale rule that uses --scale-rule-tcp-concurrency
- D. A memory scale rule with a Utilization value of 70 percent
C is correct.
Explanation: TCP scaling adjusts the replica count based on concurrent TCP connections. Use it for WebSocket servers, gRPC services, and other workloads with long-lived connections. TCP scaling supports scale-to-zero. When all connections close, the app can scale to zero after the cool-down period.
A is incorrect: HTTP scaling counts requests received over a 15-second window. It fits short-lived request-response traffic, not persistent connections.
B is incorrect: CPU scale rules can't scale an app to zero. At least one replica must run so that utilization can be measured.
D is incorrect: Memory scale rules also require at least one running replica, so they can't meet the scale-to-zero requirement.
Q008 - Question
Your AI application has 400 registered users. During the busiest interval, about 120 users are online, and about 15 of them run code analysis at the same time. Each active conversation uses its own code interpreter session. You must limit cost and protect downstream resources. How should you configure the session pool's maximum concurrent sessions?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Set the maximum near 15, and then use load testing to verify whether requests wait or fail at that limit.
- B. Set the maximum to 400 so that every registered user can have a session.
- C. Set the maximum to 120 so that every online user can have a session.
- D. Increase the cooldown period instead of setting a maximum concurrent sessions value.
A is correct.
Explanation: Base the maximum concurrent sessions value on expected concurrent active requests, not on total or online users. The limit also sets a boundary for cost and downstream resource use. Load testing shows how requests behave when the pool reaches that boundary.
B is incorrect: Total registered users overstate concurrency and remove the cost and resource protection that the limit provides.
C is incorrect: Online users aren't all running code. A value of 120 overestimates demand when only about 15 conversations execute code at once.
D is incorrect: Cooldown controls how long an idle session keeps its state. It doesn't limit how many sessions run at the same time, and a longer cooldown holds capacity longer.
Q009 - Question
After an update, upstream systems report that your AI API has intermittent delays and missing responses. You already know the name of the revision that serves traffic. You suspect that the revision scales to zero or has a crash loop. You need to verify the running instances of that specific revision. Which command should you run?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. az containerapp registry list -n ai-api -g rg-aca-demo
- B. az containerapp replica list -n ai-api -g rg-aca-demo --revision MyRevision
- C. az containerapp env show --name aca-env-demo --resource-group rg-aca-demo
- D. az containerapp revision list -n ai-api -g rg-aca-demo --all
B is correct.
Explanation: Replicas are the running instances of a revision. Use az containerapp replica list with --revision to check whether the specified revision runs and scales as expected. This check helps you detect scale-to-zero or crash loops during an incident.
A is incorrect: The registry list shows registries configured on the app. Use it to debug image pull failures, not to check running instances.
C is incorrect: The env show command shows environment details. Use it to confirm region and configuration, not replica state.
D is incorrect: The revision list shows active and inactive revisions. It confirms which version is active, but it doesn't show the running instances of a revision.
Q010 - Question
An upstream dependency for your AI background worker in Azure Container Apps has an outage. The failure isn't limited to one revision. While you mitigate the incident, you must ensure that the container app doesn't start any replicas. What should you do?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Run az containerapp stop for the container app.
- B. Rely on the existing scale-to-zero configuration for the container app.
- C. Run az containerapp restart for the container app.
- D. Run az containerapp revision deactivate for the newest revision.
A is correct.
Explanation: Stopping is an explicit app-level action. Use it when you need to ensure that the app doesn't start replicas while you mitigate an incident.
B is incorrect: Scale to zero depends on your scaling configuration, so it doesn't ensure that replicas stay stopped during the incident.
C is incorrect: Restart forces replicas to restart and can clear transient failure states, but replicas run again after the restart.
D is incorrect: Deactivation removes one revision from traffic. The problem affects the whole app, so other revisions can still run.
Q011 - Question
Your team stores a container app definition in a YAML file in source control and reviews all configuration changes like code. A developer needs to move the ai-api app to the image myregistry.azurecr.io/ai-api:v2. The change must stay consistent with the reviewed configuration. What should the developer do?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Update the image value in the YAML file, submit it for review, and apply it with az containerapp update --yaml.
- B. Run az containerapp update --yaml ./containerapp.yml --image myregistry.azurecr.io/ai-api:v2 so the --image flag overrides the YAML file.
- C. Run az containerapp up with the new image so the command creates a new environment and supporting resources.
- D. Delete the app and recreate it with az containerapp create --image myregistry.azurecr.io/ai-api:v2 without a YAML file.
A is correct.
Explanation: When you use --yaml, the YAML file is the source of truth. Change the image in the YAML file, review the change, and apply it with az containerapp update --yaml. The update creates a new revision that you can validate.
B is incorrect: When you use --yaml, other flags are ignored. The --image flag doesn't override the YAML file.
C is incorrect: Use az containerapp up for prototypes or a first deployment. It doesn't keep the change in the reviewed YAML file.
D is incorrect: Recreating the app without the YAML file bypasses code review and can cause configuration drift across environments.
Q012 - Question
Many container apps in your AI solution pull production images from Azure Container Registry. Rotating registry credentials across these apps takes significant effort. You need a predictable authentication method that reduces the number of secrets your team must rotate and follows least privilege. What should you do?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Run az containerapp registry set with --username and --password for each app, and rotate the password on a schedule.
- B. Omit credentials when you run az containerapp registry set, and let Azure CLI infer them for Azure Container Registry.
- C. Assign a managed identity to each app and grant the identity broad permissions on the registry so that all image operations succeed.
- D. Assign a managed identity to each app, grant the identity the AcrPull role on the registry, and run az containerapp registry set with --identity system.
D is correct.
Explanation: For Azure Container Registry, managed identity avoids long-lived credentials in deployment scripts. Grant the identity only the AcrPull role, and then configure the registry with --identity. Assign the identity and grant pull permissions before you use this pattern.
A is incorrect: Username and password works broadly, but it increases secret management overhead across many apps.
B is incorrect: Azure CLI can sometimes infer credentials, but this behavior isn't predictable. Use an explicit authentication approach in production.
C is incorrect: Least privilege requires that you assign only AcrPull to identities that pull images.
Q013 - Question
You deploy an AI inference API to Azure Kubernetes Service (AKS). The Pods stay in the Pending status and never transition to Running. You need to identify the cause and resolve the issue. What should you do?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Run kubectl describe pod to check for insufficient memory or CPU, and review resource requests or add nodes to the cluster.
- B. Run kubectl logs to check for application startup errors, and verify that all environment variables are configured.
- C. Run kubectl describe svc to check for empty endpoints, and verify that the Pod labels match the Service selector.
- D. Run az acr repository list to confirm the image exists, and verify the image path in the Deployment manifest.
A is correct.
Explanation: Pending Pods usually indicate that Kubernetes cannot schedule the Pods due to insufficient resources. You should run kubectl describe pod to check for insufficient memory or CPU events, and then review resource requests or scale the cluster.
B is incorrect: Checking application logs is the solution for CrashLoopBackOff, which occurs when the application exits or crashes on startup.
C is incorrect: Checking for empty endpoints is the solution when a Service has no endpoints because the selector does not match the Pod labels.
D is incorrect: Verifying the image path and registry is the solution for ImagePullBackOff, which occurs when Kubernetes cannot pull the container image.
Q014 - Question
Your team runs a model inference API and a background data enrichment worker on Azure Kubernetes Service (AKS). When slowdowns occur, the team wants a quick visual assessment first. Then the team wants direct access to cluster resources, logs, and events for in-depth analysis. Which approach should you use?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Use only kubectl commands for all inspection, because the Azure portal doesn't show workload logs.
- B. Use only the Azure portal Diagnose and solve problems feature, because it replaces log and event inspection.
- C. Use Azure portal tools such as the Workloads blade and Live Logs for visual inspection, and use kubectl commands for detailed investigation.
- D. Wait for users to report errors, and then redeploy both components to reset their state.
C is correct.
Explanation: The evidence states that the Azure portal offers visual tools like the Workloads blade, Live Logs, and the Diagnose and solve problems feature for quick inspection. For detailed investigation, kubectl commands give you direct access to cluster resources, logs, and events. Many developers use both approaches together.
A is incorrect: The Azure portal provides Live Logs to view container output, so it is not limited to non-log inspection.
B is incorrect: Diagnose and solve problems provides guided troubleshooting, but you still use kubectl for detailed investigation of resources, logs, and events.
D is incorrect: The goal of monitoring is to spot issues before they affect users. Redeploying does not identify whether the cause is application code, Kubernetes configuration, or the cluster.
Q015 - Question
A team runs a containerized Python FastAPI inference API on AKS as a single Pod that they created directly. They want Kubernetes to maintain a set number of copies of the API and restart any copy that crashes. What should they configure?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. A Service of type ClusterIP that selects the Pod by label
- B. A Deployment that specifies the container image and the number of replicas
- C. A scheduled task that runs kubectl logs against the Pod
- D. A Service of type ExternalName that maps to the Pod
B is correct.
Explanation: A Deployment defines how many replicas of your application run and specifies the container image and configuration. If a Pod crashes, the Deployment restarts it. You typically don't run single Pods directly because Deployments manage Pods for you.
A is incorrect: A ClusterIP Service provides a stable internal endpoint for Pods. It doesn't control how many Pods run or restart crashed Pods.
C is incorrect: kubectl logs shows log output for troubleshooting. It doesn't create or restart Pods.
D is incorrect: An ExternalName Service maps a Service to an external DNS name. It doesn't manage Pod replicas.
Q016 - Question
You deploy an AI inference API to Azure Kubernetes Service (AKS). The API needs environment-specific endpoint settings and API keys for upstream services. It also needs a location for user uploads that keeps data when Pods restart. You must not hardcode values into the container image. Which approach should you use?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Store the endpoint settings and API keys in a ConfigMap, and store user uploads in the container filesystem.
- B. Store the endpoint settings in a Secret, store the API keys in a ConfigMap, and store user uploads in a PersistentVolumeClaim (PVC).
- C. Store the endpoint settings in a ConfigMap, store the API keys in a Secret, and store user uploads in a PersistentVolumeClaim (PVC).
- D. Store the endpoint settings, API keys, and user uploads in a single PersistentVolumeClaim (PVC).
C is correct.
Explanation: Use ConfigMaps to separate nonsensitive configuration, such as endpoints, from code. Use Secrets to protect sensitive values, such as API keys. Use a PVC for durable storage that survives Pod restarts.
A is incorrect: API keys are sensitive values and belong in a Secret. Also, data in the container filesystem is lost when a Pod restarts.
B is incorrect: This option reverses the roles. API keys are sensitive and belong in a Secret, and endpoint settings are nonsensitive and belong in a ConfigMap.
D is incorrect: A PVC provides durable storage for data. It does not externalize settings or protect sensitive values the way ConfigMaps and Secrets do.
Q017 - Question
A production AI workload on AKS has strict compliance requirements. Secrets must remain exclusively in a dedicated vault. You need fine-grained audit trails for each secret access. Rotated secrets must reach the containers without Pod restarts. What should you use?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. The Azure Key Vault provider for Secrets Store CSI Driver to mount secrets from Azure Key Vault into Pod filesystems.
- B. Azure App Configuration with Key Vault references, resolved by the App Configuration Kubernetes Provider into Kubernetes Secret resources.
- C. Opaque Kubernetes Secrets created with kubectl create secret generic and referenced with secretKeyRef.
- D. Secret manifests with literal values stored in a GitOps repository and applied with kubectl apply.
A is correct.
Explanation: The Secrets Store CSI Driver mounts secrets from Key Vault into Pod filesystems as files. Secrets remain in Key Vault. The driver polls for changes and updates mounted secrets without Pod restarts. Choose direct Key Vault integration when you need fine-grained audit trails for each secret access.
B is incorrect: The App Configuration Kubernetes Provider generates native Kubernetes Secret resources with the resolved values. Use this option when you want to manage secret mappings centrally, not when secrets must remain exclusively in the vault.
C is incorrect: Kubernetes Secrets store the values in the cluster. Environment variable values from secretKeyRef are set at Pod start, so you must trigger a Deployment rollout after you update them.
D is incorrect: Never commit Secret manifests with literal values to source control. This option also does not keep secrets in a dedicated vault.
Q018 - Question
You run kubectl apply for an inference API Deployment that references myregistry.azurecr.io/inference-api:v1.0. The Pods show the status ImagePullBackOff. What should you do to diagnose and resolve the issue?
Domain: Develop containerized solutions on Azure (20–25%) Type: Single choice
- A. Increase the memory and CPU requests in the Deployment manifest.
- B. Run az aks scale to add more nodes to the cluster.
- C. Run kubectl describe pod to review the events, then verify the image path and tag and confirm the image exists by using az acr repository list.
- D. Update the Service selector so that it matches the Pod labels.
C is correct.
Explanation: ImagePullBackOff means Kubernetes can't pull your container image from the registry. Run kubectl describe pod and review the events section for the pull error. Then verify the image path, tag, and registry in your manifest, and confirm the image was built and pushed with az acr repository list.
A is incorrect: Resource requests affect scheduling. Requests that are too large cause Pending Pods, not image pull failures.
B is incorrect: Adding nodes with az aks scale resolves exhausted cluster capacity for Pending Pods. More nodes can't pull an image that doesn't exist or has the wrong path.
D is incorrect: A selector mismatch causes a Service to have no endpoints. It doesn't prevent the image from being pulled.
Q019 - Question
You build an AI-powered recommendation engine on Azure Cosmos DB for NoSQL. You use the default indexing policy. A new feature requires queries that sort products by more than one property with ORDER BY. What should you do to support this query pattern?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Change the container partition key to the property used in the ORDER BY clause.
- B. Switch the container from manual throughput to autoscale throughput.
- C. Customize the indexing policy and add a composite index for the ORDER BY pattern.
- D. Move the products to a new database that uses shared throughput.
C is correct.
Explanation: The default indexing policy automatically indexes the properties in your items, which reduces upfront index setup. As query patterns change, you might still need to customize the indexing policy. For example, some ORDER BY patterns require composite indexes.
A is incorrect: The partition key controls how data is distributed across physical partitions. It does not define index paths. You also cannot change a container's partition key without creating a new container and migrating data.
B is incorrect: Autoscale changes how RU/s scale with demand. It does not add the indexes that some ORDER BY patterns require.
D is incorrect: Shared throughput at the database level changes how capacity is allocated across containers. It does not change the indexing policy of a container.
Q020 - Question
Support agents use a semantic search application built on Azure Cosmos DB for NoSQL. Agents report that searches such as "error code 0x80070005 access denied" return conceptually related articles but often miss articles that contain the exact error code. You need results that reflect both exact term matches and semantic similarity in a single ranked list. What should you configure?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Add a WHERE clause that returns only documents with a VectorDistance score greater than 0.7.
- B. Use ORDER BY RANK RRF to combine VectorDistance on titleEmbedding and VectorDistance on contentEmbedding.
- C. Change the vector policy dataType to float16 and increase the TOP N value to 100.
- D. Add a full-text policy and a full-text index on /content, and use ORDER BY RANK RRF with VectorDistance and FullTextScore.
D is correct.
Explanation: Hybrid search combines vector similarity with full-text scoring by using the RRF function. Configure a full-text policy and a full-text index on the text property, then rank with VectorDistance and FullTextScore so that the error code is matched exactly and "access denied" is matched semantically.
A is incorrect: A similarity threshold filters out low-quality semantic matches but doesn't add exact keyword matching for terms such as error codes.
B is incorrect: Multi-vector search combines scores from multiple embedding spaces. Both scores are still semantic, so exact term matching isn't added.
C is incorrect: float16 reduces storage, and a larger TOP N increases RU consumption. Neither change adds keyword matching.
Q021 - Question
A data ingestion pipeline writes large volumes of document chunks and embeddings to Azure Cosmos DB for NoSQL in real time. A reporting job queries the container once per day. The indexing policy includes /* and defines 12 composite indexes. Write RU consumption exceeds the budget. You need to reduce overall RU cost for this workload. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Add composite indexes for every combination of properties that the reporting job might use.
- B. Include only the properties that the reporting job filters or sorts on, and remove composite indexes that the job doesn't need.
- C. Change the account default consistency level to strong.
- D. Keep /* in includedPaths and add range indexes on the embedding arrays.
B is correct.
Explanation: This workload is write-heavy. Every index increases write latency and RU consumption because indexes are updated synchronously during writes. Minimize indexed properties and composite indexes. When queries are infrequent, an occasional higher query cost can be acceptable compared to constantly elevated write costs.
A is incorrect: Each composite index adds write overhead. Creating composite indexes for edge cases increases write costs without proportional read benefits.
C is incorrect: Write RU costs don't vary by consistency level at the operation level. Strong consistency doubles read RU cost and increases write latency in multi-region accounts.
D is incorrect: Embedding arrays in range indexes increase storage and write overhead without query benefits. Each array element is indexed separately, and vector searches use vector indexes.
Q022 - Question
Your team has an existing Azure Cosmos DB for NoSQL account. An existing container stores support documents but was created without a vector policy. You need to store embeddings and run vector similarity queries on the support documents. Which two actions should you take? Each correct answer presents part of the solution.
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Multi-select
- A. Update the vector embedding policy on the existing container to add the /embedding path.
- B. Enable the Vector Search for NoSQL API feature on the account, for example by using az cosmosdb update --capabilities EnableNoSQLVectorSearch.
- C. Add the /embedding/* path to includedPaths in the existing container's indexing policy.
- D. Create a new container and specify the vector embedding policy and the indexing policy at creation time.
B and D are correct.
Explanation: You must enable vector search on the Azure Cosmos DB account before you use it. Vector policies must be set when the container is created, and vector search is only supported on new containers, so create a new container with both policies.
A is incorrect: Vector policies can't be modified after the container exists.
C is incorrect: Adding the embedding path to a range index doesn't configure vector search. Embedding arrays don't benefit from range indexes and increase write costs.
Q023 - Question
You are designing an Azure Cosmos DB for NoSQL container for a knowledge base. Each document stores a 3,072-dimension embedding from the text-embedding-3-large model. The container will grow to millions of vectors, and similarity queries must return results quickly with high accuracy. Which vector index type should you configure on the embedding path?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. flat
- B. quantizedFlat
- C. diskANN
- D. No vector index. Include /embedding/* in the range index includedPaths.
C is correct.
Explanation: Use diskANN because it supports up to 4,096 dimensions and provides the best performance for large datasets with millions of vectors while maintaining high accuracy.
A is incorrect: The flat index performs exact brute-force search and limits vectors to 505 dimensions, so it can't index 3,072-dimension embeddings.
B is incorrect: quantizedFlat supports up to 4,096 dimensions but still performs brute-force search and is recommended for datasets up to approximately 50,000 vectors per physical partition.
D is incorrect: Embedding arrays don't benefit from standard range indexes. Including them wastes storage and increases write costs, and queries without a vector index perform a full scan.
Q024 - Question
A production RAG application queries an Azure Cosmos DB for NoSQL container that uses a DiskANN vector index. Before release, your team wants to validate how closely the indexed results match exact results for a set of test queries. Production queries must keep low latency and RU consumption. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. In the evaluation queries only, pass true as the third parameter to VectorDistance to force brute-force search, and compare the results with the indexed results.
- B. Pass true as the third parameter to VectorDistance in all production queries so that every query returns exact results.
- C. Remove the TOP N clause from the evaluation queries so that DiskANN returns exact results.
- D. Change the distanceFunction in the vector policy to euclidean so that VectorDistance returns exact results.
A is correct.
Explanation: Indexed search trades a small amount of accuracy for better performance. Use brute-force search for scenarios that require exact accuracy, such as evaluation and testing, and keep indexed search for production queries.
B is incorrect: Brute-force search compares the query vector against every document, consumes significantly more RUs, has higher latency, and scales poorly as data grows.
C is incorrect: Removing TOP N doesn't force exact search. Without TOP N, the query attempts to return all documents, which consumes excessive RUs and causes high latency.
D is incorrect: The distance function determines how similarity is calculated, not whether search is exact. Vector policies also can't be modified after the container exists.
Q025 - Question
A nightly Python job uses psycopg to load about 50,000 archived agent messages into the messages table in Azure Database for PostgreSQL. The job currently runs one INSERT statement per row in a loop and takes too long. You need the highest load throughput. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Use executemany to send the rows as parameterized INSERT statements.
- B. Use cur.copy() with a COPY messages (conversation_id, role, content) FROM STDIN statement and write each row with copy.write_row().
- C. Build one multi-row INSERT statement by concatenating the message values into the SQL string.
- D. Keep the per-row INSERT loop and open a new connection for each row so that connections stay short-lived.
B is correct.
Explanation: For larger datasets of 10,000 rows or more, the COPY command provides the highest performance. It is often two to 10 times faster than individual inserts. In psycopg, use cur.copy() and write each record with copy.write_row().
A is incorrect: Use executemany for inserting hundreds to a few thousand rows. For 50,000 rows, COPY provides higher throughput.
C is incorrect: Never use string formatting or concatenation to build queries with data values. Use parameterized queries to prevent SQL injection.
D is incorrect: Creating new database connections is expensive because each requires network handshakes, authentication, and server-side resource allocation. A new connection per row makes the load slower.
Q026 - Question
You design a products table for filtered vector search in Azure Database for PostgreSQL. Most queries filter by category_id and a price range before they order by vector similarity. Products also have vendor-specific attributes that differ between products and are rarely filtered. Which two design choices should you make? Each correct answer presents part of the solution. Select two.
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Multi-select
- A. Store price in a JSONB column and create a GIN index to accelerate BETWEEN range filters.
- B. Store category_id and price as structured columns and create a composite B-tree index on (category_id, price).
- C. Partition the two-million-row table by price range to speed up the category filters.
- D. Store the vendor-specific, rarely filtered attributes in a JSONB column.
B and D are correct.
Explanation: Structured columns with a B-tree index on (category_id, price) let PostgreSQL narrow the candidate rows before it calculates vector distances. A JSONB column gives you schema flexibility for dynamic attributes that you rarely filter on. This combination optimizes the common case and keeps the design flexible.
A is incorrect: GIN indexes support containment queries but not range queries. A price range filter on JSONB requires an expression index or a sequential scan.
C is incorrect: Consider partitioning when tables exceed tens of millions of rows and queries filter by the partition key. Here, the queries filter mainly by category, so partitioning by price adds complexity and doesn't match the query pattern.
Q027 - Question
You store 1536-dimensional dense embeddings from text-embedding-ada-002 in a vector(1536) column in Azure Database for PostgreSQL. Storage costs are increasing as the document table grows. You must reduce embedding storage and keep acceptable similarity search quality. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Change the column to sparsevec(1536) so that only nonzero values are stored.
- B. Change the column to vector(768) and keep the same embedding model.
- C. Change the column to halfvec(1536) after benchmarking confirms that search quality is acceptable.
- D. Create an IVFFlat index on the column to compress the stored vectors.
C is correct.
Explanation: The halfvec type stores elements as 16-bit floating-point numbers and uses half the storage of the vector type. Use halfvec only after you benchmark and verify that half precision provides acceptable search quality for your use case.
A is incorrect: The sparsevec type is for models that produce sparse embeddings where most elements are zero. Dense embeddings from text-embedding-ada-002 are not sparse.
B is incorrect: The vector(n) dimension must match the output dimension of the embedding model. A 768-dimension column causes insertion errors for 1536-dimension embeddings.
D is incorrect: An index does not reduce the size of the vector column. An IVFFlat index adds about 1 to 1.5 times the vector column size in storage.
Q028 - Question
Attorneys use a vector-only search over legal documents in Azure Database for PostgreSQL. When they search for a specific case name such as "Smith v. Jones," documents that contain that exact name sometimes do not appear in the results. You must return exact term matches and keep semantic relevance. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Lower the cosine distance threshold from 0.5 to 0.3.
- B. Combine vector similarity with PostgreSQL full-text search by using ts_rank in a weighted hybrid score, and create a GIN index for keyword matching.
- C. Average the query embedding with embeddings of example documents before you search.
- D. Change the distance operator from <=> to <->.
B is correct.
Explanation: Hybrid search combines vector similarity with keyword-based full-text search to capture both semantic and lexical relevance. Use it when users search for names, codes, or terms that must match exactly. A GIN index on the full-text search vector speeds up keyword matching.
A is incorrect: A stricter distance threshold increases precision, but it still ranks only by semantic distance. It can remove more results and does not add exact term matching.
C is incorrect: Averaging vectors is for finding documents similar to several related examples. It does not add lexical matching for an exact case name.
D is incorrect: L2 distance is still a vector distance metric. Changing the operator does not add keyword matching, and most text embedding models are optimized for cosine similarity.
Q029 - Question
Your legal search application stores 1536-dimensional embeddings in an embedding column. You plan to move to a new embedding model that outputs 3,072 dimensions. The application must keep serving consistent results during the migration, and you must be able to roll back if the new model underperforms. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Update the existing embedding column in place in batches with the new model's vectors.
- B. Store new-model vectors in the existing embedding column and add a column that tags each row with the model version.
- C. Regenerate the embeddings with the new model and then run REINDEX on the existing index.
- D. Add an embedding_v2 vector(3072) column with its own HNSW index, backfill it in batches, validate results, and then switch queries to the new column.
D is correct.
Explanation: A parallel column strategy lets you populate new embeddings alongside the existing ones, compare search quality, and switch over when you are ready. This approach adds temporary storage overhead, but it avoids downtime and lets you roll back if the new model underperforms.
A is incorrect: Overwriting embeddings in place returns inconsistent results while the migration is in progress, and you lose the old vectors needed for rollback.
B is incorrect: You can't mix embeddings from different models in the same column. Different models produce vectors with different dimensions and semantic relationships.
C is incorrect: Rebuilding the index does not solve the dimension change or the need to keep the old embeddings available during validation and rollback.
Q030 - Question
Your team built a proof of concept for an AI agent on an Azure Database for PostgreSQL server that uses the Burstable tier. You are preparing the solution for production. The production workload has steady, predictable resource requirements, and the agent makes many short-lived database calls to store individual messages. You must use the built-in PgBouncer connection pooler. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Keep the Burstable tier and set the pgbouncer.enabled server parameter to true.
- B. Keep the Burstable tier and change the application connection string to use port 6432.
- C. Change the server to the General Purpose tier, enable PgBouncer, and connect on port 6432.
- D. Change the server to the Memory Optimized tier, enable PgBouncer, and continue to connect on port 5432.
C is correct.
Explanation: Built-in PgBouncer is available only on the General Purpose and Memory Optimized tiers. General Purpose fits production workloads with steady, predictable resource requirements. After you enable PgBouncer, connect on port 6432 instead of 5432. You can change compute tiers after deployment with a brief restart.
A is incorrect: The Burstable tier doesn't support the built-in PgBouncer feature, so you can't enable it on this tier.
B is incorrect: Changing the port doesn't enable pooling. PgBouncer isn't available on the Burstable tier, so port 6432 doesn't route through a pooler.
D is incorrect: Port 5432 is the standard PostgreSQL port for direct connections. To use PgBouncer, connect on port 6432. Memory Optimized also targets large in-memory workloads, which this scenario doesn't require.
Q031 - Question
A retailer generates image embeddings from a vision model for two million product photos. The embeddings aren't pre-normalized. The similarity search must return results in under 10 ms, and the team wants to keep Redis memory costs low. Which vector field configuration should you use?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. TYPE FLOAT64, DISTANCE_METRIC L2, algorithm HNSW
- B. TYPE FLOAT32, DISTANCE_METRIC COSINE, algorithm FLAT
- C. TYPE FLOAT32, DISTANCE_METRIC IP, algorithm HNSW
- D. TYPE FLOAT32, DISTANCE_METRIC L2, algorithm HNSW
D is correct.
Explanation: Use FLOAT32 because it uses 4 bytes per dimension and provides enough precision for AI embeddings. Use L2 for image embeddings and spatial data. Use HNSW for datasets over 10,000 vectors when you need queries under 10 ms and can accept 95-99% accuracy.
A is incorrect: FLOAT64 uses 8 bytes per dimension, which doubles memory usage and slows distance calculations without a meaningful accuracy gain for embeddings.
B is incorrect: COSINE is the metric for text embeddings, and FLAT compares the query against every vector. FLAT query time grows linearly, so it doesn't meet the latency target at two million vectors.
C is incorrect: Use IP only with pre-normalized embeddings or specialized models that require it. These embeddings aren't pre-normalized.
Q032 - Question
Your team prepares a workstation for a lab task. The task creates an Azure Managed Redis resource by using the Azure CLI. It then runs a Python app that loads, stores, and searches vector data with metadata. Which preparation meets the documented prerequisites for this task?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Python 3.10, the latest version of the Azure CLI, and Visual Studio Code
- B. An Azure subscription with permission to create an Azure Managed Redis instance with an enterprise SKU, Python 3.12 or greater, the latest version of the Azure CLI, and the Azure CLI redisenterprise extension
- C. An Azure subscription with permission to create an Azure Managed Redis instance with an enterprise SKU, Python 3.12 or greater, and the latest version of the Azure CLI without extensions
- D. An Azure subscription with read-only access to existing resources, Python 3.12 or greater, the latest version of the Azure CLI, and the Azure CLI redisenterprise extension
B is correct.
Explanation: To create the resource and run the app, you need an Azure subscription with permission to create an Azure Managed Redis instance with an enterprise SKU. You also need Visual Studio Code, Python 3.12 or greater, the latest version of the Azure CLI, and the Azure CLI redisenterprise extension.
A is incorrect: Python 3.10 doesn't meet the Python 3.12 or greater requirement, and this option omits the redisenterprise extension and the subscription permission.
C is incorrect: The Azure CLI alone isn't enough. You must also install the redisenterprise extension.
D is incorrect: Read-only access doesn't let you create the Azure Managed Redis instance. You need permission to create an instance with an enterprise SKU.
Q033 - Question
You develop a Python AI application that uses the redis-py library. The Azure Managed Redis instance uses the OSS clustering policy. You need to configure the client connection correctly for this clustering policy. What should you use?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. redis.Redis with port 6380
- B. redis.Redis with decode_responses=False
- C. redis.Redis with a Microsoft Entra ID credential_provider
- D. redis.cluster.RedisCluster
D is correct.
Explanation: The clustering policy chosen for an Azure Managed Redis instance impacts the connection method. If you are working with an instance using the OSS clustering policy, you need to use redis.cluster.RedisCluster instead of redis.Redis for your connection.
A is incorrect: Port 6380 is the default for Azure Cache for Redis, not Azure Managed Redis (which uses 10000). Changing the port does not configure the client for the OSS clustering policy.
B is incorrect: The decode_responses=False parameter configures the client to work with raw bytes instead of strings. It does not configure the client for the OSS clustering policy.
C is incorrect: A credential_provider configures Microsoft Entra ID authentication. It does not configure the client for the OSS clustering policy.
Q034 - Question
You design vector storage for a product catalog in Azure Managed Redis. Each product has nested specifications, several variant records, and two embeddings: one for the description and one for the main image. The application already exchanges product data in JSON format. Which storage approach should you use?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Store each product as a Redis JSON document with embeddings converted by tolist(), and create the index with IndexType.JSON and JSONPath field names such as $.embedding.
- B. Store each product as a Redis Hash with both embeddings converted by tobytes(), and create the index with IndexType.HASH.
- C. Store each product as a Redis Hash, and serialize the nested specifications and variants into a single text field.
- D. Store each product as a Redis JSON document with embeddings converted by tobytes(), and create the index with IndexType.HASH.
A is correct.
Explanation: Choose Redis JSON when your data has nested structures, you need multiple vectors per item, or your application already uses JSON. Convert each embedding to a numeric array with tolist(), and define the index with IndexType.JSON, $. JSONPath field names, and the as_name parameter for query references.
B is incorrect: Redis Hash stores flat field-value pairs only. It's the better choice for simple records when you need maximum memory efficiency and query speed, but it doesn't support nested objects.
C is incorrect: Flattening nested data into one text field works against the documented guidance. Hash is intended for data models that won't need nested objects.
D is incorrect: JSON documents store vectors as numeric arrays, not binary bytes, and a JSON-based index must use IndexType.JSON.
Q035 - Question
A web application for an AI shopping assistant stores shopping carts, user preferences, and authentication tokens in browser cookies. Bandwidth usage and latency increase as session data grows. You need to reduce request overhead and keep fast access to session data by using Azure Managed Redis. What should you do?
Domain: Develop AI solutions by using Azure data management services (25–30%) Type: Single choice
- A. Store only a session identifier in the cookie, and use that key to retrieve the full session data from Azure Managed Redis.
- B. Keep the full session data in the cookie, and copy it to Azure Managed Redis as a backup.
- C. Store the session data in the backend database, and cache only static headers and footers in Azure Managed Redis.
- D. Compress the session data and keep it in the cookie so that no server-side session store is needed.
A is correct.
Explanation: Storing too much information directly in cookies negatively impacts performance because each cookie is transmitted with every HTTP request and response. The typical solution stores only a session identifier in the cookie, then uses that key to retrieve full session data from Azure Managed Redis, which delivers sub-millisecond latency.
B is incorrect: Keeping the full session data in the cookie means every request still transmits the data, which does not reduce bandwidth usage or latency.
C is incorrect: Caching static headers and footers is a content cache pattern, not a session store pattern. Querying session data from a backend database is slower than retrieving it from Azure Managed Redis.
D is incorrect: Compressing the session data in the cookie still transmits the data with every request and response, and this approach does not use Azure Managed Redis as a session store.
Q036 - Question
Your image classification pipeline uses a Service Bus Standard tier namespace. Each request includes an image of about 2 MB. The business doesn't want to upgrade to the Premium tier only for message size. The security team also wants access to the image data controlled separately from access to the messaging layer. What should you do?
Domain: Connect to and consume Azure services (20–25%) Type: Single choice
- A. Upload each image to Azure Blob Storage and send a Service Bus message that contains the blob URI and request metadata.
- B. Group the image messages into a ServiceBusMessageBatch so that the batch handles the payload size.
- C. Serialize each image as base64 inside a JSON body and set the content_type property to application/json.
- D. Set a longer time_to_live value on each message so that large messages have more time to transfer.
A is correct.
Explanation: Use the claim-check pattern. The Standard tier supports messages up to 256 KB. With claim-check, the message carries only a reference, so it works within any tier's size limits. You can also scope Blob Storage access independently from Service Bus access.
B is incorrect: Batching reduces network round trips. The SDK keeps a batch within the maximum message size for your tier, so batching doesn't let a single 2-MB message fit.
C is incorrect: The content_type property only describes the encoding format. A base64-encoded 2-MB image still exceeds the 256-KB Standard tier limit.
D is incorrect: The time_to_live property controls how long a message stays in the queue before it expires. It has no effect on the maximum message size.
Q037 - Question
An orchestration processes a large claim archive with one flat fan-out/fan-in pattern. Measurements show that the fan-in step is slow because one orchestrator loads and aggregates every individual document result. You need to reduce the state that one orchestration handles. You also want an independent status for each part of the archive. What should you do?
Domain: Connect to and consume Azure services (20–25%) Type: Single choice
- A. Replace context.task_all() with sequential awaits so that the orchestrator aggregates one result at a time.
- B. Return the extracted text from each activity so that the orchestrator can aggregate results without extra storage reads.
- C. Increase the number of activities that the orchestrator schedules at once in the single flat fan-out.
- D. Partition the archive by policy or month. Use sub-orchestrations that return counts and result locations, and have the parent aggregate those summaries.
D is correct.
Explanation: One orchestrator performs the fan-in step on one worker at a time, so a flat fan-in can become a bottleneck. Sub-orchestrations let child orchestrations process separate partitions and return compact summaries. The parent aggregates those summaries instead of every document result. Each partition also gets an independent status.
A is incorrect: Sequential processing adds the latency of every item and doesn't reduce the results that one orchestrator must aggregate.
B is incorrect: Returning extracted text expands orchestration history, and Durable Functions must load that data during replay.
C is incorrect: Scheduling more work at once doesn't reduce the aggregation work in one orchestrator. A sudden activity burst can also exceed model rate limits or saturate a downstream data service.
Q038 - Question
You maintain a Durable Functions app that processes insurance claim documents. The orchestrator passes full extracted document text to activities. Activities return complete model responses to the orchestrator. Replay time and storage operations are increasing. The orchestration history also retains sensitive claim content. What should you change?
Domain: Connect to and consume Azure services (20–25%) Type: Single choice
- A. Pass the full document text as orchestration input instead of activity input so that activities receive smaller payloads.
- B. Store the content in Azure Blob Storage, and pass a blob URL with compact metadata such as the document ID and confidence score.
- C. Compress the complete model responses inside the orchestrator before the orchestrator returns its result.
- D. Add a storage access token to the orchestration input so that each activity can read the content directly.
B is correct.
Explanation: Durable Functions serializes orchestration inputs, activity inputs, and activity outputs into the durable store. Large payloads increase storage operations, replay time, memory use, and the amount of sensitive data retained in history. Store large content in a data service such as Azure Blob Storage. Pass only a reference and the fields that the orchestrator needs for its next decision.
A is incorrect: Orchestration inputs are also serialized into orchestration history, so the full text stays in the durable store.
C is incorrect: The orchestrator still receives and handles the complete model responses. Keep large results out of the orchestration, write them from the activity, and return a reference.
D is incorrect: Don't place secrets or access tokens in orchestration inputs or outputs. Use managed identities and a service such as Azure Key Vault for credentials.
Q039 - Question
A Service Bus-triggered function calls Azure AI Document Intelligence to extract text from each document. Monitoring shows extra latency on every invocation because the function creates a new DocumentIntelligenceClient and credential each time it runs. You need to reduce this per-invocation overhead. What should you do?
Domain: Connect to and consume Azure services (20–25%) Type: Single choice
- A. Replace the SDK client with a Document Intelligence output binding in function_app.py.
- B. Initialize the credential and the DocumentIntelligenceClient at the module level, outside the function handler.
- C. Define the Document Intelligence client settings in host.json so that the runtime creates the client.
- D. Set maxConcurrentCalls to 1 in host.json so that each message has dedicated resources.
B is correct.
Explanation: You should initialize SDK clients outside the function handler at the module level. The objects persist across invocations on the same instance, so you avoid repeated connection setup, configuration loading, and token acquisition on every function invocation.
A is incorrect: Azure AI Document Intelligence does not have a dedicated binding. You must create SDK clients directly in your function code for this service.
C is incorrect: The host.json file contains global runtime settings such as timeouts, logging, and trigger concurrency. It does not create SDK clients for your code.
D is incorrect: The maxConcurrentCalls setting controls how many messages each instance processes at the same time. It does not remove the cost of creating a client on every invocation.
Q040 - Question
A function app has a system-assigned managed identity. One function writes classification results by using a Cosmos DB output binding configured with CosmosDBConnection__accountEndpoint. The same function calls Azure OpenAI Service through an SDK client that uses DefaultAzureCredential. You need to remove all connection strings and API keys and apply least privilege. Which two role assignments should you grant to the managed identity? Each correct answer presents part of the solution.
Domain: Connect to and consume Azure services (20–25%) Type: Multiple choice
- A. Key Vault Secrets User
- B. Cosmos DB Built-in Data Contributor
- C. Cognitive Services OpenAI User
- D. Azure Service Bus Data Sender
B and C are correct.
Explanation: An identity-based Cosmos DB connection uses the account endpoint, and the managed identity needs the Cosmos DB Built-in Data Contributor role to write documents. Azure OpenAI Service has no dedicated binding, so the SDK client authenticates with DefaultAzureCredential. In production, this credential uses the managed identity, which needs the Cognitive Services OpenAI User role.
A is incorrect: The Key Vault Secrets User role is required for Key Vault references. This design removes stored secrets, so the function does not read from Key Vault.
D is incorrect: The Azure Service Bus Data Sender role is for Service Bus output bindings. This function does not write to Service Bus, so the role adds access the function does not need.
Q041 - Question
Your team runs a monitoring tool that checks whether the secrets that a RAG pipeline requires exist in Azure Key Vault. The tool must read secret names and properties. It must not read secret values. Which built-in role should you assign to the tool's identity?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Key Vault Secrets User
- B. Key Vault Reader
- C. Key Vault Contributor
- D. Key Vault Secrets Officer
B is correct.
Explanation: Key Vault Reader grants read access to vault metadata, such as secret names and properties. It doesn't reveal secret values or key material. Use it for monitoring and discovery tools that verify which secrets exist.
A is incorrect: Key Vault Secrets User grants read access to secret values. The tool must not read values, so this role grants more access than the tool requires.
C is incorrect: Key Vault Contributor is a control plane role. It manages the vault resource, but it doesn't grant data plane operations on secrets, keys, or certificates.
D is incorrect: Key Vault Secrets Officer grants full management permissions on secrets, including create, update, and delete. Assign it to operators or CI/CD pipelines that manage the secret lifecycle, not to a read-only monitoring tool.
Q042 - Question
You add error handling to a Python service that calls get_secret() on a SecretClient. You want the service to retry only failures that might resolve without operator intervention. Missing secrets and missing RBAC role assignments must fail immediately so that operators can diagnose them. Which exception type should trigger a retry?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. ResourceNotFoundError
- B. HttpResponseError
- C. ServiceRequestError
- D. Any exception that get_secret() raises
C is correct.
Explanation: ServiceRequestError indicates a network-level problem, such as a DNS resolution failure or a timeout. This type of failure might resolve on retry.
A is incorrect: ResourceNotFoundError means that the secret name doesn't exist in the vault. This is a configuration error that requires operator intervention, so a retry won't resolve it.
B is incorrect: HttpResponseError covers authentication and authorization failures, such as an identity that lacks the required RBAC role. Retrying doesn't fix a missing role assignment.
D is incorrect: Retrying every exception also retries configuration and authorization errors. These errors won't resolve on their own, and retries delay diagnosis.
Q043 - Question
Your AI pipeline needs three settings: Pipeline:BatchSize, OpenAI:DeploymentName, and Storage:AccountKey. The storage account key must rotate on a schedule, and you need an audit trail of each access. The app must retrieve all three settings through a single load() call. Which design should you use?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Store all three settings as direct values in App Configuration, and use labels to separate environments.
- B. Store Pipeline:BatchSize and OpenAI:DeploymentName as direct values in App Configuration, and store the account key in Key Vault with a Key Vault reference in App Configuration.
- C. Store all three settings in Key Vault, and read them with the Key Vault SDK.
- D. Store the account key in both App Configuration and Key Vault as independent values, so the app can read either copy.
B is correct.
Explanation: Values that grant access to a resource belong in Key Vault, which provides rotation, expiration policies, and per-object audit logging. Nonsensitive values such as batch sizes and deployment names belong in App Configuration. A Key Vault reference keeps App Configuration as the single entry point, so one load() call returns all three settings.
A is incorrect: App Configuration doesn't provide per-object audit logging, HSM-backed encryption, expiration policies, or automated rotation for the account key.
C is incorrect: Key Vault has stricter throttling limits and doesn't support labels, feature flags, or snapshots for nonsensitive settings. This design also doesn't use a single load() call.
D is incorrect: Independent copies in two services can drift apart, and the app might read a stale value. A Key Vault reference stores only a pointer, so there's one source of truth.
Q044 - Question
A Python app uses a managed identity and calls load() with both credential and keyvault_credential set to DefaultAzureCredential. The identity has the App Configuration Data Reader role on the store. Regular settings load, but the provider throws an error that names the Key Vault and the secret for OpenAI:ApiKey. What should you do to resolve the error?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Assign the App Configuration Data Owner role on the App Configuration store to the managed identity.
- B. Copy the secret value from Key Vault into App Configuration as a regular key-value pair.
- C. Append a specific version identifier to the secret identifier URI in the Key Vault reference.
- D. Assign the Key Vault Secrets User role on the referenced Key Vault to the managed identity.
D is correct.
Explanation: To resolve Key Vault references, the app identity needs a role on each service: App Configuration Data Reader on the store and Key Vault Secrets User on each referenced vault. When Key Vault access is missing, the provider throws an error that identifies the vault and secret it couldn't access.
A is incorrect: The identity already reads the store. App Configuration Data Owner adds write access to the store but doesn't grant access to secrets in Key Vault.
B is incorrect: Storing the secret directly in App Configuration removes Key Vault protections such as audit logging, rotation, and HSM-backed encryption.
C is incorrect: Pinning a version doesn't fix a missing permission, and it stops the app from picking up rotated secrets without a configuration change.
Q045 - Question
Operators often change several related settings in Azure App Configuration at the same time, such as Pipeline:BatchSize and Pipeline:RetryCount. The Python app must use the new values without a restart, and operators must control when the changes take effect so that the app doesn't apply a partial update. What should you implement?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Configure load() with refresh_on set to WatchKey("Sentinel") and a refresh_interval, call config.refresh() in the request handler, and update the Sentinel key after you change the other settings.
- B. Configure load() with a refresh_interval only, and rely on the provider to push updated values to the app without calling refresh().
- C. Call load() once at startup and restart the app after operators finish changing the settings.
- D. Create a new label for each set of changes, and redeploy the app with a new label_filter value in its SettingSelector.
A is correct.
Explanation: With the sentinel key pattern, operators change the settings first and then update the Sentinel key. When the provider detects the change to the watched key, it reloads the entire configuration, so the app receives all changes together on the next refresh cycle.
B is incorrect: The provider doesn't push values. You must call refresh() explicitly, and you need refresh_on to define the sentinel key to watch.
C is incorrect: A restart doesn't meet the requirement to pick up changes without restarting the app.
D is incorrect: Redeploying the app for each change removes the benefit of updating configuration without redeployment.
Q046 - Question
You publish an Azure dashboard that contains pinned log query tiles from the pipeline's Application Insights resource. A team member has read permission on the dashboard resource. When the team member opens the dashboard, the log query tiles display an access error. What should you do?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Select Share and then Publish again to republish the dashboard.
- B. Add a Markdown tile that lists the on-call contacts and runbook links.
- C. Re-pin each log query with a render operator so the chart type is preserved.
- D. Grant the team member read permission on the Application Insights resource.
D is correct.
Explanation: Dashboard access uses Azure role-based access control (RBAC). Team members need read permissions on both the dashboard resource and the underlying data sources. Without access to the Application Insights resource, tiles that query that resource display an access error.
A is incorrect: Publishing makes the dashboard available to other users, but it doesn't grant access to the underlying Application Insights data.
B is incorrect: A Markdown tile adds context to the dashboard. It doesn't change permissions on the data source.
C is incorrect: The render operator preserves the visualization when you pin results. It doesn't resolve a permissions error.
Q047 - Question
Four Python services in a RAG pipeline use the Azure Monitor OpenTelemetry Distro and send telemetry to the same Application Insights resource. On the Application Map, all four services appear as a single node. You need each service to appear as a separate node while keeping the services logically grouped. What should you configure?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Set a unique service.name resource attribute for each service and use the same service.namespace resource attribute for all four services.
- B. Set the same service.name resource attribute for all four services and a unique service.instance.id resource attribute for each service.
- C. Pass a different connection_string value to configure_azure_monitor() in each service.
- D. Add a span attribute named service.name to each custom span by using set_attribute().
A is correct.
Explanation: Application Insights forms the cloud role name from service.namespace combined with service.name. A unique service.name per service creates a separate Application Map node for each service. A shared service.namespace groups the services together logically.
B is incorrect: service.instance.id distinguishes multiple instances of the same service. With the same service.name, the services still share one cloud role name.
C is incorrect: The connection string tells the exporter which Application Insights resource receives telemetry. It does not set the cloud role name, and the services need to send to the same resource.
D is incorrect: Span attributes describe a single operation and appear in customDimensions. The cloud role name comes from resource attributes, which apply to all telemetry from a service.
Q048 - Question
Your client requires two monitoring outcomes for the content moderation pipeline. First, the team must be notified when the 95th-percentile response time for any service exceeds three seconds. Second, the team must be notified about gradual performance degradation that might never exceed a fixed threshold. No one watches the monitoring views continuously. Which two solutions should you implement? Each correct answer presents part of the solution.
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Multiple choice
- A. A metric alert rule on average server response time with a threshold of three seconds.
- B. A log search alert rule that runs a KQL query returning services where percentile(duration, 95) exceeds 3000.
- C. A dashboard tile that renders p50, p95, and p99 latency as a time chart.
- D. Smart detection performance anomaly detection in Application Insights.
B and D are correct.
Explanation: A log search alert can use any KQL query, including percentile calculations, so it enforces the specific three-second p95 threshold. Smart detection learns the normal behavior of your application and detects performance that is worse than its historical norm, including gradual degradation that never exceeds a fixed number.
A is incorrect: An average response time can stay within acceptable limits even when a significant percentage of requests is slow. It doesn't enforce a p95 requirement.
C is incorrect: A dashboard tile shows latency, but it requires someone to be watching. It doesn't notify the team.
Q049 - Question
You query the dependencies table for failed calls from the classification service. Most failures target the model inference endpoint and have a resultCode of 429. What should you do to address these failures?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Investigate a server-side failure in the model inference endpoint code.
- B. Adjust throughput to the model inference endpoint or implement retry policies.
- C. Investigate network issues between the classification service and the endpoint.
- D. Query the exceptions table for unhandled exceptions in the ingestion API.
B is correct.
Explanation: An HTTP 429 result code from a model inference endpoint indicates rate limiting. You address rate limiting by adjusting throughput or implementing retry policies.
A is incorrect: A server-side failure in the dependency is indicated by an HTTP 500 result code, not 429.
C is incorrect: Network issues or an overloaded service typically appear as a timeout with no result code. These failures have a result code of 429.
D is incorrect: The failed calls originate from the classification service and target the inference endpoint. The ingestion API isn't the source of these dependency failures.
Q050 - Question
You monitor a retrieval-augmented generation (RAG) pipeline that runs as four microservices. Your client requires 95th-percentile response times under three seconds. You need to detect when response times trend above this target and trigger an alert. Which observability pillar should you use for this requirement?
Domain: Secure, monitor, and troubleshoot Azure solutions (20–25%) Type: Single choice
- A. Distributed tracing
- B. Logs
- C. Metrics
- D. Context propagation
C is correct.
Explanation: Metrics provide aggregate numerical measurements over time, such as request counts, error rates, and response-time percentiles. Use metrics to detect trends and set alerting thresholds for service-level objectives, such as a 95th-percentile target.
A is incorrect: Distributed tracing shows the path and timing of individual requests. Use traces to find where a problem occurs after metrics show that something changed.
B is incorrect: Logs are timestamped records of discrete events within a service. Logs explain why a specific operation behaved a certain way, but they do not provide aggregate percentile measurements for alert thresholds.
D is incorrect: Context propagation carries trace and span IDs across service boundaries. It is a mechanism that connects spans into one trace, not an observability pillar for percentile alerting.