Scalable Real-Time Anomaly Detection on Market Data Streams with Kinesis and Managed Flink Random Cut Forest

Answer Correct answer: A — Kinesis ingests the market data and Managed Flink scales automatically, with built-in RANDOM_CUT_FOREST detecting stream anomalies and no clusters to run.

A financial company receives a high volume of real-time market data streams from an external provider. The streams consist of thousands of JSON records every second. The company needs to implement a scalable solution on AWS to identify anomalous data points. Which solution will meet these requirements with the LEAST operational overhead?

  1. Ingest real-time data into Amazon Kinesis data streams. Use the built-in RANDOM_CUT_FOREST function in Amazon Managed Service for Apache Flink to process the data streams and to detect data anomalies. Correct Answer
  2. Ingest real-time data into Amazon Kinesis data streams. Deploy an Amazon SageMaker endpoint for real-time outlier detection. Create an AWS Lambda function to detect anomalies. Use the data streams to invoke the Lambda function.
  3. Ingest real-time data into Apache Kafka on Amazon EC2 instances. Deploy an Amazon SageMaker endpoint for real-time outlier detection. Create an AWS Lambda function to detect anomalies. Use the data streams to invoke the Lambda function.
  4. Send real-time data to an Amazon Simple Queue Service (Amazon SQS) FIFO queue. Create an AWS Lambda function to consume the queue messages. Program the Lambda function to start an AWS Glue extract, transform, and load (ETL) job for batch processing and anomaly detection.

Community Votes

A
100%

100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Kinesis Data Streams provides the managed real-time ingestion and Amazon Managed Service for Apache Flink supplies a managed, scalable stream processor with a built-in RANDOM_CUT_FOREST function for anomaly detection, so no clusters or self-managed Kafka have to be operated.

A financial company receives thousands of JSON records per second from an external market data provider and needs a scalable AWS solution to identify anomalous data points with the least operational overhead. The volume is high and continuous, so the solution must be a managed real-time stream processing path rather than self-managed infrastructure.

Standing up Apache Kafka on EC2 instances, which adds self-managed cluster operations to a fully managed streaming problem. Another common error is assuming a SageMaker real-time outlier detection endpoint plus Lambda is simpler, when it adds a model and function to maintain for a capability Flink provides natively.

Community Discussion (3 comments)

dduenas 👍 1 Selected: B
Tricky question. RANDOM_CUT_FOREST is not available as a function in flink. it is a legacy function of Kinesis Data Analytics SQL. there is a RandomCutForestOperator, but is different that the mentioned function. ( https://aws.amazon.com/blogs/big-data/real-time-anomaly-detection-via-random-cut-forest-in-amazon-managed-service-for-apache-flink/)
Saransundar 👍 3 Selected: A
Option A High-volume real-time: Kinesis Data Streams Scalable: Managed Apache Flink Anomaly detection: RANDOM_CUT_FOREST Low overhead: Fully managed services
GiorgioGss 👍 3 Selected: A
https://docs.aws.amazon.com/kinesisanalytics/latest/sqlref/sqlrf-random-cut-forest.html "Detects anomalies in your data stream."

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The workload is a high-volume real-time stream where thousands of records arrive every second and must be scanned for anomalies, and the requirement is the least operational overhead. Kinesis Data Streams is a managed service for real-time ingestion at that scale, and Amazon Managed Service for Apache Flink is a managed, automatically scaling stream processor, so neither requires the company to run clusters. Crucially, Flink provides the RANDOM_CUT_FOREST function as a built-in for streaming anomaly detection, which detects anomalies directly in the stream without provisioning a separate model endpoint. The vote was 86 for A, the only option that scores meaningfully. Saransundar mapped each part of the requirement to its service, and GiorgioGss cited the Kinesis Data Analytics SQL reference for RANDOM_CUT_FOREST, quoting that it detects anomalies in your data stream.

Why the Other Options Are Wrong

Deploying a SageMaker endpoint for real-time outlier detection with a Lambda function invoked by the data streams (B) works functionally but adds a model to train, host, and monitor plus a function to maintain, when the managed Flink path already includes the detection capability, so it carries more operational overhead. Running Apache Kafka on Amazon EC2 instances with the same SageMaker and Lambda detection (C) compounds that overhead by adding a self-managed Kafka cluster that the company must provision, patch, and operate, which directly contradicts the least-overhead requirement. Sending data to an SQS FIFO queue and having Lambda start a Glue ETL job for batch processing and anomaly detection (D) converts a real-time requirement into a batch one, and a FIFO queue adds ordering constraints and polling latency that make it unsuited to a high-throughput stream.

Community Comment Notes

The community was heavily in favor at 86 votes for A, and the one substantive dissent came from dduenas, who called it a tricky question and argued that RANDOM_CUT_FOREST is a legacy function of Kinesis Data Analytics SQL rather than a Flink function, noting that Flink has a RandomCutForestOperator that is a different thing, and linked the AWS blog on real-time anomaly detection via Random Cut Forest in Managed Flink. The majority position, held by Saransundar and GiorgioGss, treats the RANDOM_CUT_FOREST capability as built in, which is the reading the question intends. The dissent is worth noting because it questions the exact naming of the function rather than the architecture.

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide