Increase the ALB target group health check grace period for the ECS service
A company runs a website by using an Amazon Elastic Container Service (Amazon ECS) service that is connected to an Application Load Balancer (ALB). The service was in a steady state with tasks responding to requests successfully. A DevOps engineer updated the task definition with a new container image and deployed the new task definition to the service. The DevOps engineer noticed that the service is frequently stopping and starting new tasks because the ALB healtth checks are failing. What should the DevOps engineer do to troubleshoot the failed deployment?
Community Votes
100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The symptom is a start-then-stop cycle driven by health check failures on newly deployed tasks, which points at startup timing rather than networking or capacity, and the health check grace period is the control that governs exactly that window (B). eugene2owl made the same observation, that the engineer only updated the image in the task definition so a security group change could not be involved, and that a new container image plausibly needs longer to become healthy. Options A, C, and D either address problems that were ruled out by the fact that only the image changed, or make the deployment worse by requiring more capacity or by checking more frequently during the window when the container is not yet ready.
The service was stable until a new task definition with a new container image was deployed, after which tasks repeatedly start and are then stopped because the ALB health checks fail. This pattern indicates the new tasks are being killed before they finish starting, which is what the health check grace period controls: the number of seconds after a target enters InService during which the load balancer does not send health check requests, giving a slow-starting container time to begin responding. Increasing the grace period on the service's ALB target group lets the new tasks pass startup before health checks can mark them unhealthy.
Ensuring a security group associated with the service allows traffic from the ALB (A) — as eugene2owl noted, the engineer only updated the container image in the task definition, so a security group change could not have caused this; and the service was in a steady state before, which shows the security groups already permitted ALB traffic. Increasing the service minimum healthy percent (C) — this changes how many tasks must be healthy during a deployment, which forces the service to keep launching replacement tasks and makes the rolling deployment slower and more likely to exhaust capacity; it does not give a slow-starting container more time to become healthy. Decreasing the ALB health check interval (D) — DKM and the other commenters identified the grace period as the correct lever, while a shorter interval means health checks are attempted sooner and more often, which makes it more likely that a still-starting task is marked unhealthy rather than less likely.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The service was in a steady state with tasks responding successfully, and the problem began only after a new task definition with a new container image was deployed. The tasks now start and are then stopped because ALB health checks fail, which is the signature of tasks being terminated before they finish starting rather than of a networking or capacity fault. The control that governs this window is the ALB target group's health check grace period, which is the number of seconds after a target enters the InService state during which the load balancer does not send health check requests, allowing a container that starts slowly to begin responding before it can be judged unhealthy. Increasing the grace period for the service's target group therefore lets the newly deployed tasks complete startup and pass health checks instead of being cycled (B). Srikantha explained exactly this, that newly deployed tasks may take time to become healthy and respond to ALB health checks, and DKM supplied the corresponding elbv2 modify-target-group-attributes command setting health_check.grace_period. B is the correct answer.Why the Other Options Are Wrong
A ensures that a security group associated with the service allows traffic from the ALB. As eugene2owl pointed out, the engineer only updated the container image in the task definition, so no security group was modified; moreover the service was previously in a steady state with tasks responding successfully, which demonstrates that the security groups already permitted traffic from the load balancer. C increases the service minimum healthy percent setting. This governs how many tasks in a service must be in the healthy state for the deployment to be considered successful, so raising it forces the service to keep launching replacement tasks and makes the deployment slower and more likely to run out of capacity; it does not extend the time a slow-starting container has to become healthy and therefore does not stop the start-stop cycle. D decreases the ALB health check interval. A shorter interval causes the load balancer to attempt health checks sooner and more frequently, which increases the likelihood that a task that has not finished starting is marked unhealthy and killed, worsening the very behaviour described. B is correct.Community Comment Notes
Community voted B unanimously. Srikantha explained that after deploying a new version of the ECS service, tasks may take time to become healthy and respond to ALB health checks, which is what the grace period accommodates. DKM supplied the exact API call, elbv2 modify-target-group-attributes with Key=health_check.grace_period, confirming that this is the setting to change. eugene2owl gave the reasoning that rules out option A, that the engineer only updated the image so he could not have affected the security group, and that a new Docker image is the plausible reason for slower startup. Ky_24 and awsarchitect5 both selected B. No alternative received support.Official Reference
Related Analysis
Practice All DOP-C02 Questions
Access 85 questions with complete answers and detailed explanations.
View Full DOP-C02 Practice Test →