How to Stream Google Cloud Logs to an On-Premises SIEM

Cloud Logging & Data Integration
Answer Correct answer: B — Route logs to a Pub/Sub topic via a logging sink and process them with a primary Dataflow pipeline while configuring a secondary pipeline to replay failed messages.

You work for a multinational organization that has systems deployed across multiple cloud providers, including Google Cloud. Your organization maintains an extensive on-premises security information and event management (SIEM) system. New security compliance regulations require that relevant Google Cloud logs be integrated seamlessly with the existing SIEM to provide a unified view of security events. You need to implement a solution that exports Google Cloud logs to your on-premises SIEM by using a push-based, near real-time approach. You must prioritize fault tolerance, security, and auto scaling capabilities. In particular, you must ensure that if a log delivery fails, logs are re-sent. What should you do?

  1. Create a Pub/Sub topic for log aggregation. Write a custom Python script on a Cloud Function Leverage the Cloud Logging API to periodically pull logs from Google Cloud and forward the logs to the SIEM. Schedule the Cloud Function to run twice per day.
  2. Collect all logs into an organization-level aggregated log sink and send the logs to a Pub/Sub topic. Implement a primary Dataflow pipeline that consumes logs from this Pub/Sub topic and delivers the logs to the SIEM. Implement a secondary Dataflow pipeline that replays failed messages. Correct Answer
  3. Deploy a Cloud Logging sink with a filter that routes all logs directly to a syslog endpoint. The endpoint is based on a single Compute Engine hosted on Google Cloud that routes all logs to the on-premises SIEM. Implement a Cloud Function that triggers a retry action in case of failure.
  4. Utilize custom firewall rules to allow your SIEM to directly query Google Cloud logs. Implement a Cloud Function that notifies the SIEM of a failed delivery and triggers a retry action.

Community Votes

B
100%

100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Tests knowledge of managed streaming services versus manual or polling approaches, with the trap being options that suggest single VMs, scheduled pulls, or direct network queries instead of scalable message queues and stream processors.

This question tests the optimal architecture for securely and reliably streaming Google Cloud logs to an on-premises SIEM in near real-time. The correct solution leverages Pub/Sub and Dataflow to ensure fault tolerance, auto-scaling, and automatic retry capabilities.

Option C is frequently selected by candidates who overlook the explicit requirement for auto-scaling, failing to realize that a single Compute Engine instance creates a bottleneck and lacks the necessary resilience.

Community Discussion (5 comments)

Zek 👍 1 Selected: B
B - https://cloud.google.com/architecture/stream-logs-from-google-cloud-to-splunk
MoAk 👍 1 Selected: B
B 100%.
KLei 👍 1 Selected: B
use pub/sub. A is wrong as it says that "periodically pull logs" - Not near real-time and need programing works.
BondleB 👍 1 Selected: B
https://cloud.google.com/architecture/stream-logs-from-google-cloud-to-splunk
yokoyan 👍 1 Selected: B
I think it's B.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option B correctly implements a push-based, near real-time architecture by routing logs through a Pub/Sub topic, which acts as a durable, auto-scaling buffer. A Dataflow pipeline then consumes these messages and streams them to the on-premises SIEM, providing built-in auto-scaling and managed error handling. By configuring a secondary pipeline or dead-letter queue to replay failed messages, the solution satisfies the strict fault tolerance and retry requirements without manual intervention. This pattern aligns perfectly with Google Cloud’s official guidance for integrating cloud telemetry with enterprise SIEMs.

Why the Other Options Are Wrong

Option A relies on a scheduled pull mechanism running twice daily, which directly contradicts the near real-time and push-based requirements. Option C provisions a single Compute Engine instance, creating a single point of failure and eliminating the required auto-scaling capability for handling variable log volumes. Option D suggests opening firewall rules for direct querying, which bypasses the controlled export pipeline, introduces significant security risks, and does not implement a reliable push-based delivery mechanism with automatic retries.

Community Comment Notes

Candidates consistently validated this approach by referencing Google’s official architecture documentation for streaming logs to third-party platforms like Splunk. Several learners noted that the polling schedule in option A explicitly disqualifies it for real-time use cases. Others emphasized that leveraging Pub/Sub as an intermediate buffer is the standard practice for decoupling log generation from downstream processing. As one contributor highlighted, relying on message queues prevents custom code bottlenecks during traffic spikes.

Official Reference

Exam Strategy

When designing cloud-to-on-premises data pipelines, always prioritize managed, fully hosted services like Pub/Sub and Dataflow over custom scripts or single VMs. Verify that every architectural choice explicitly maps to non-functional requirements such as near real-time latency, auto-scaling, and automated retry mechanisms before selecting an option.

Frequently Asked Questions

Why is a single Compute Engine instance unsuitable for log forwarding?

It creates a single point of failure and cannot auto-scale to handle traffic spikes, violating the fault tolerance and scaling requirements.

Can Cloud Functions replace Dataflow for this log export task?

Cloud Functions lack the built-in state management and high-throughput stream processing needed for reliable, near real-time log delivery at scale.

Related Analysis

← Back to PCSE Study Guide