AWS Glue Job vs Crawler for SQL Server Migration
A company is migrating its database servers from Amazon EC2 instances that run Microsoft SQL Server to Amazon RDS for Microsoft SQL Server DB instances. The company's analytics team must export large data elements every day until the migration is complete. The data elements are the result of SQL joins across multiple tables. The data must be in Apache Parquet format. The analytics team must store the data in Amazon S3. Which solution will meet these requirements in the MOST operationally efficient way?
Community Votes
53% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The core concept is minimizing manual overhead in ETL pipelines; the trap is assuming a Crawler is needed for metadata discovery when the source structure (view) is already defined and stable.
This question evaluates the operational efficiency of using AWS Glue Jobs versus Crawlers for migrating SQL Server data to Parquet format in S3. It establishes that direct JDBC extraction via Glue Job is superior because schema is known.
Candidates often choose Option C, believing that a Glue Crawler is mandatory to catalog the database schema before running a job. This adds unnecessary operational complexity when the schema is static.
Community Discussion (13 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option A is the most operationally efficient solution because it leverages AWS Glue's managed ETL capabilities directly against the source database view. Since the view schema is fixed and known, there is no need for a Glue Crawler to discover or update metadata. The Glue Job can connect via JDBC, extract the data, transform it to Apache Parquet, and load it into S3 in a single streamlined step.Why the Other Options Are Wrong
Option C includes an unnecessary Glue Crawler step, adding latency and management overhead without providing value since the schema is predefined. Option B involves multiple steps: writing CSV from SQL Agent, triggering Lambda, and converting formats, which is inefficient and fragile. Option D requires managing a custom Lambda function with JDBC drivers and connection pooling logic, increasing maintenance burden compared to the managed Glue service.Community Comment Notes
The community was split between A and C. Some users argued for C, citing decoupling or metadata management needs. However, many commenters correctly pointed out that crawlers are redundant for known schemas. As user mtrianac noted, "using an AWS Glue Crawler is not necessary... Glue can directly connect to the database... without additional steps." Another user, leonardoFelipe, questioned if crawlers understand views, reinforcing the preference for direct job execution.Exam Strategy
When designing ETL solutions, always look for the option that minimizes manual configuration steps. If the data schema is known and stable (like a SQL View), skip the metadata discovery phase (Crawler) and go straight to the transformation and loading phase (Job).
Frequently Asked Questions
Why is a Glue Crawler not needed for SQL Server views?
Glue Crawlers are used to discover schema changes in data stores like S3 or dynamic databases. For a static SQL View with a known, unchanging schema, the Glue Job can define the schema directly in its script or JDBC connector, making the crawler redundant.
Can AWS Glue read directly from SQL Server?
Yes, AWS Glue supports JDBC connections to various relational databases, including Amazon RDS for SQL Server and EC2-hosted SQL Server instances, allowing direct data extraction without intermediate file storage.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →