Replace an idle EC2 crawling fleet with Lambda and write results to S3
A company is running a web-crawling process on a list of target URLs to obtain training documents for machine learning training algorithms. A fleet of Amazon EC2 t2.micro instances pulls the target URLs from an Amazon Simple Queue Service (Amazon SQS) queue. The instances then write the result of the crawling algorithm as a .csv file to an Amazon Elastic File System (Amazon EFS) volume. The EFS volume is mounted on all instances of the fleet. A separate system adds the URLs to the SQS queue at infrequent rates. The instances crawl each URL in 10 seconds or less. Metrics indicate that some instances are idle when no URLs are in the SQS queue. A solutions architect needs to redesign the architecture to optimize costs. Which combination of steps will meet these requirements MOST cost-effectively? (Choose two.)
Community Votes
100% of anonymous learners picked answer BE. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The workload is queue-driven and bursty with idle gaps, which is the pattern Lambda is built for because billing follows invocations, so there is no cost when the queue is empty, and the .csv output is a file that belongs in object storage rather than a shared file system.
A fleet of t2.micro instances pulls URLs from an SQS queue, crawls each in ten seconds or less, and writes .csv results to a shared EFS volume. URLs arrive at infrequent rates, so instances sit idle when the queue is empty, and the architecture must be redesigned for cost.
Moving to larger instance types and scaling the fleet in. The problem is idle capacity, not insufficient capacity per instance, so bigger instances and a smaller fleet raise per-crawl cost while the idle-time cost remains. Choosing a database such as Neptune or Aurora Serverless for the output is also wrong because the requirement is simply to store .csv files.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The decisive observation is that instances are idle when the SQS queue has no URLs, so the fleet is paying for capacity it does not use. Converting the crawler into an AWS Lambda function with an SQS event source means the function is invoked only when messages are available and the company pays only for the compute time consumed, which is why B is correct. For the results, each crawl produces a.csv file, and Amazon S3 is the natural destination because it stores objects of any size, requires no running server, and costs a fraction of the provisioned capacity a fleet or a database would require while the queue is empty, which is why E is correct.Why the Other Options Are Wrong
A: The issue is idle capacity rather than throughput, so moving to m5.8xlarge instances and halving the fleet increases the cost per unit of work while still paying for idle instances, and it does not change the underlying paying-for-idle-time problem. C: Amazon Neptune is a graph database designed for relationship queries and must run in a VPC, and the crawl output is flat.csv files with no graph structure, so it is both wrong and expensive. D: Amazon Aurora Serverless is a relational database, and storing.csv results in a database adds per-request query cost and provisioned capacity for data that is simply files to be retained and read later.Community Comment Notes
The community voted 100 to 0 for B and E, and the top-voted comment gave the reasoning directly, Lambda plus S3, with the crawler not needing a database at all. Another commenter described option A as scaling out when what is needed is scaling to zero, and noted that the database options are expensive and unsuitable for this output shape.Official Reference
Related Analysis
Practice All SAP-C02 Questions
Access 85 questions with complete answers and detailed explanations.
View Full SAP-C02 Practice Test →