I am creating an ECS EC2 type cluster and deploying my container images in that using ECS service. Deployment is succesfull and all the required resources are created and container is deployed onto the ECS cluster.
But when i am trying to destroy the resources, the operation gets stuck at ECS service destroy and gets failed after 20 min timeout (Althuogh it deletes the ECS service in the backend but remains in destroying state) and when i do terraform destroy again it deletes the remaining resources.
Please provide your suggestion if anyone have faced this kind of issue.
I have attached the screenshot of the timeout error also for reference.
the service can be only deleted if you scale it down to 0 replicas. It’s not possible to remove services if you have one or more replicas (desired state). It’s the stupid ECS thing.
during the stuck phase i can see in the portal that service is deleted from ECS cluster and showing the status as ‘Draining’. So It’s ECS who is not updating the actual status.
I’m experiencing same issue. And I ended up being delete terraform.tf and remove all deployed resources manually. Are there better ways to handle this kind issue?
provisioner “local-exec” {
when = destroy
command = <<EOF
echo “Update service desired count to 0 before destroy.” #Get region out of cluster
REGION=(echo {self.cluster} | cut -d’:’ -f4)
echo “Region: $REGION” #Set the Service desired count to 0
aws ecs update-service --region REGION --cluster {self.cluster} --service ${self.name} --desired-count 0 --force-new-deployment
echo “Update service command executed successfully.”
EOF
}
timeouts { #create = “10m” # Timeout for resource creation
delete = “5m” # Timeout for resource deletion
}
}
The ECS service destruction process gets stuck and eventually fails after a 20-minute timeout, even though the ECS service is deleted in the backend. Strangely, upon running terraform destroy again, the remaining resources are successfully removed. I have attached a screenshot of the timeout error for reference.
Had the exact same symptom: destroy stuck on aws_ecs_service for 10+ minutes, service showing DRAINING with 0 running tasks. Running destroy a second time always succeeded immediately.
The root cause was agentConnected: false on the container instance. Terraform was destroying the route table association in parallel with the ECS service. This removed the subnet’s route to the internet gateway mid-drain, cutting the ECS agent’s outbound connection to the ECS control plane. With the agent disconnected, ECS couldn’t confirm the task had stopped, so the service sat in DRAINING indefinitely.
The fix was adding aws_route_table_association.public to depends_on on the aws_ecs_service resource: