Replace an idle EC2 crawling fleet with Lambda and write results to S3

Answer Correct answers: B, E — Convert the crawler to a Lambda function pulling from SQS and write the .csv results to Amazon S3.

A company is running a web-crawling process on a list of target URLs to obtain training documents for machine learning training algorithms. A fleet of Amazon EC2 t2.micro instances pulls the target URLs from an Amazon Simple Queue Service (Amazon SQS) queue. The instances then write the result of the crawling algorithm as a .csv file to an Amazon Elastic File System (Amazon EFS) volume. The EFS volume is mounted on all instances of the fleet. A separate system adds the URLs to the SQS queue at infrequent rates. The instances crawl each URL in 10 seconds or less. Metrics indicate that some instances are idle when no URLs are in the SQS queue. A solutions architect needs to redesign the architecture to optimize costs. Which combination of steps will meet these requirements MOST cost-effectively? (Choose two.)

  1. Use m5.8xlarge instances instead of t2.micro instances for the web-crawling process. Reduce the number of instances in the fleet by 50%.
  2. Convert the web-crawling process into an AWS Lambda function. Configure the Lambda function to pull URLs from the SQS queue. Correct Answer
  3. Modify the web-crawling process to store results in Amazon Neptune.
  4. Modify the web-crawling process to store results in an Amazon Aurora Serverless MySQL instance.
  5. Modify the web-crawling process to store results in Amazon S3. Correct Answer

Community Votes

BE
100%

100% of anonymous learners picked answer BE. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The workload is queue-driven and bursty with idle gaps, which is the pattern Lambda is built for because billing follows invocations, so there is no cost when the queue is empty, and the .csv output is a file that belongs in object storage rather than a shared file system.

A fleet of t2.micro instances pulls URLs from an SQS queue, crawls each in ten seconds or less, and writes .csv results to a shared EFS volume. URLs arrive at infrequent rates, so instances sit idle when the queue is empty, and the architecture must be redesigned for cost.

Moving to larger instance types and scaling the fleet in. The problem is idle capacity, not insufficient capacity per instance, so bigger instances and a smaller fleet raise per-crawl cost while the idle-time cost remains. Choosing a database such as Neptune or Aurora Serverless for the output is also wrong because the requirement is simply to store .csv files.

Community Discussion (4 comments)

AzureDP900 👍 1
Option B involves converting the web-crawling process into an AWS Lambda function, which will allow the company to pay only for the compute time consumed by the function. This approach can help reduce costs compared to running EC2 instances. Converting the web-crawling process to a Lambda function also eliminates the need to maintain and monitor a fleet of EC2 instances, reducing operational costs. Option E involves modifying the web-crawling process to store results in Amazon S3, which is an object storage service that can store large amounts of data. This approach can help reduce costs compared to storing data on EFS or other block-based file systems. Storing results in S3 also provides high availability and durability for the data, meeting the requirements of a production-grade system.
pangchn 👍 3 Selected: BE
BE lamda + S3 the process don't need a database
Dgix 👍 3 Selected: BE
A is utter rubbish - scaling out is not what we need B is optimal in terms of cost C and D involve fairly expensive databases not suitable for this use case. Moreover, Neptune must run in a VPC. E is optimal in terms of accessibility and cost
CMMC 👍 1 Selected: BE
use lambda instead of a fleet of EC2, and store the results into cost-effective S3

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The decisive observation is that instances are idle when the SQS queue has no URLs, so the fleet is paying for capacity it does not use. Converting the crawler into an AWS Lambda function with an SQS event source means the function is invoked only when messages are available and the company pays only for the compute time consumed, which is why B is correct. For the results, each crawl produces a.csv file, and Amazon S3 is the natural destination because it stores objects of any size, requires no running server, and costs a fraction of the provisioned capacity a fleet or a database would require while the queue is empty, which is why E is correct.

Why the Other Options Are Wrong

A: The issue is idle capacity rather than throughput, so moving to m5.8xlarge instances and halving the fleet increases the cost per unit of work while still paying for idle instances, and it does not change the underlying paying-for-idle-time problem. C: Amazon Neptune is a graph database designed for relationship queries and must run in a VPC, and the crawl output is flat.csv files with no graph structure, so it is both wrong and expensive. D: Amazon Aurora Serverless is a relational database, and storing.csv results in a database adds per-request query cost and provisioned capacity for data that is simply files to be retained and read later.

Community Comment Notes

The community voted 100 to 0 for B and E, and the top-voted comment gave the reasoning directly, Lambda plus S3, with the crawler not needing a database at all. Another commenter described option A as scaling out when what is needed is scaling to zero, and noted that the database options are expensive and unsuitable for this output shape.

Official Reference

Related Analysis

Practice All SAP-C02 Questions

Access 85 questions with complete answers and detailed explanations.

View Full SAP-C02 Practice Test →

← Back to SAP-C02 Study Guide