AWS Glue Job vs Crawler for SQL Server Migration

Answer Correct answer: A — Create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create an AWS Glue job that selects the data directly from the view and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day.

A company is migrating its database servers from Amazon EC2 instances that run Microsoft SQL Server to Amazon RDS for Microsoft SQL Server DB instances. The company's analytics team must export large data elements every day until the migration is complete. The data elements are the result of SQL joins across multiple tables. The data must be in Apache Parquet format. The analytics team must store the data in Amazon S3. Which solution will meet these requirements in the MOST operationally efficient way?

  1. Create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create an AWS Glue job that selects the data directly from the view and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day. Correct Answer
  2. Schedule SQL Server Agent to run a daily SQL query that selects the desired data elements from the EC2 instance-based SQL Server databases. Configure the query to direct the output .csv objects to an S3 bucket. Create an S3 event that invokes an AWS Lambda function to transform the output format from .csv to Parquet.
  3. Use a SQL query to create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create and run an AWS Glue crawler to read the view. Create an AWS Glue job that retrieves the data and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day.
  4. Create an AWS Lambda function that queries the EC2 instance-based databases by using Java Database Connectivity (JDBC). Configure the Lambda function to retrieve the required data, transform the data into Parquet format, and transfer the data into an S3 bucket. Use Amazon EventBridge to schedule the Lambda function to run every day.

Community Votes

A
53%
C
47%

53% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The core concept is minimizing manual overhead in ETL pipelines; the trap is assuming a Crawler is needed for metadata discovery when the source structure (view) is already defined and stable.

This question evaluates the operational efficiency of using AWS Glue Jobs versus Crawlers for migrating SQL Server data to Parquet format in S3. It establishes that direct JDBC extraction via Glue Job is superior because schema is known.

Candidates often choose Option C, believing that a Glue Crawler is mandatory to catalog the database schema before running a job. This adds unnecessary operational complexity when the schema is static.

Community Discussion (13 comments)

taka5094 👍 7 Selected: C
Choice A) is almost the same approach, but it doesn't use the AWS Glue crawler, so have to manage the view's metadata manually.
Christina666 👍 7 Selected: C
Leveraging SQL Views: Creating a view on the source database simplifies the data extraction process and keeps your SQL logic centralized. Glue Crawler Efficiency: Using a Glue crawler to automatically discover and catalog the view's metadata reduces manual setup. Glue Job for ETL: A dedicated Glue job is well-suited for the data transformation (to Parquet) and loading into S3. Glue jobs offer built-in scheduling capabilities. Operational Efficiency: This approach minimizes custom code and leverages native AWS services for data movement and cataloging.
Eltanany 👍 1 Selected: A
I'll go with A
Certified101 👍 1 Selected: A
A is correct - no need for crawler
plutonash 👍 3 Selected: A
the scrawler is not necessary, use GLUE job to read data from sql server and transfert to S3 with Apache Parquet format is enough.
mtrianac 👍 3 Selected: A
No, in this case, using an AWS Glue Crawler is not necessary. The schema is already defined in the SQL Server database, as the created view contains the required structure (columns and data types). AWS Glue can directly connect to the database via JDBC, extract the data, transform it into Parquet format, and store it in S3 without additional steps. A crawler is useful if you're working with data that doesn't have a predefined schema (e.g., files in S3) or if you need the data to be cataloged for services like Amazon Athena. However, for this ETL flow, using just a Glue Job simplifies the process and reduces operational complexity.
michele_scar 👍 1 Selected: A
Glue crawler is useless because the schema is already in place with a SQL database
leonardoFelipe 👍 3 Selected: A
Usually, views aren't true objects in a SGBD, they're just a "nickname" for a specific query string, different of Materialized Views. So, my questions is: can glue crawler understand their metadata? I'd go with A
bakarys 👍 3 Selected: A
Option A involves creating a view in the EC2 instance-based SQL Server databases that contains the required data elements. An AWS Glue job is then created to select the data directly from the view and transfer the data in Parquet format to an S3 bucket. This job is scheduled to run every day. This approach is operationally efficient as it leverages managed services (AWS Glue) and does not require additional transformation steps. Option D involves creating an AWS Lambda function that queries the EC2 instance-based databases using JDBC. The Lambda function is configured to retrieve the required data, transform the data into Parquet format, and transfer the data into an S3 bucket. This approach could work, but managing and scheduling Lambda functions could add operational overhead compared to using managed services like AWS Glue.
GiorgioGss 👍 2 Selected: C
Just beacuse it decouples the whole architecture I will go with C
Felix_G 👍 1
Option C seems to be the most operationally efficient: It leverages Glue for both schema discovery (via the crawler) and data transfer (via the Glue job). The Glue job can directly handle the Parquet format conversion. Scheduling the Glue job ensures regular data export without manual intervention.
rralucard_ 👍 3 Selected: A
Option A (Creating a view in the EC2 instance-based SQL Server databases and creating an AWS Glue job that selects data from the view, transfers it in Parquet format to S3, and schedules the job to run every day) seems to be the most operationally efficient solution. It leverages AWS Glue’s ETL capabilities for direct data extraction and transformation, minimizes manual steps, and effectively automates the process.
evntdrvn76 👍 2
A. Create a view in the EC2 instance-based SQL Server databases that contains the required data elements. Create an AWS Glue job that selects the data directly from the view and transfers the data in Parquet format to an S3 bucket. Schedule the AWS Glue job to run every day. This solution is operationally efficient for exporting data in the required format.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option A is the most operationally efficient solution because it leverages AWS Glue's managed ETL capabilities directly against the source database view. Since the view schema is fixed and known, there is no need for a Glue Crawler to discover or update metadata. The Glue Job can connect via JDBC, extract the data, transform it to Apache Parquet, and load it into S3 in a single streamlined step.

Why the Other Options Are Wrong

Option C includes an unnecessary Glue Crawler step, adding latency and management overhead without providing value since the schema is predefined. Option B involves multiple steps: writing CSV from SQL Agent, triggering Lambda, and converting formats, which is inefficient and fragile. Option D requires managing a custom Lambda function with JDBC drivers and connection pooling logic, increasing maintenance burden compared to the managed Glue service.

Community Comment Notes

The community was split between A and C. Some users argued for C, citing decoupling or metadata management needs. However, many commenters correctly pointed out that crawlers are redundant for known schemas. As user mtrianac noted, "using an AWS Glue Crawler is not necessary... Glue can directly connect to the database... without additional steps." Another user, leonardoFelipe, questioned if crawlers understand views, reinforcing the preference for direct job execution.

Exam Strategy

When designing ETL solutions, always look for the option that minimizes manual configuration steps. If the data schema is known and stable (like a SQL View), skip the metadata discovery phase (Crawler) and go straight to the transformation and loading phase (Job).

Frequently Asked Questions

Why is a Glue Crawler not needed for SQL Server views?

Glue Crawlers are used to discover schema changes in data stores like S3 or dynamic databases. For a static SQL View with a known, unchanging schema, the Glue Job can define the schema directly in its script or JDBC connector, making the crawler redundant.

Can AWS Glue read directly from SQL Server?

Yes, AWS Glue supports JDBC connections to various relational databases, including Amazon RDS for SQL Server and EC2-hosted SQL Server instances, allowing direct data extraction without intermediate file storage.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide