Google Professional Data Engineer Exam - Topic 2 Question 22 Discussion
You are building a new application that you need to collect data from in a scalable way. Data arrives continuously from the application throughout the day, and you expect to generate approximately 150 GB of JSON data per day by the end of the year. Your requirements are:Decoupling producer from consumerSpace and cost-efficient storage of the raw ingested data, which is to be stored indefinitelyNear real-time SQL queryMaintain at least 2 years of historical data, which will be queried with SQWhich pipeline should you use to meet these requirements?
A) Create an application that provides an API. Write a tool to poll the API and write data to Cloud Storage as gzipped JSON files.
B) Create an application that writes to a Cloud SQL database to store the data. Set up periodic exports of the database to write to Cloud Storage and load into BigQuery.
C) Create an application that publishes events to Cloud Pub/Sub, and create Spark jobs on Cloud Dataproc to convert the JSON data to Avro format, stored on HDFS on Persistent Disk.
D) Create an application that publishes events to Cloud Pub/Sub, and create a Cloud Dataflow pipeline that transforms the JSON event payloads to Avro, writing the data to Cloud Storage and BigQuery.
Cruz
10 months agoShayne
10 months agoFiliberto
11 months agoDelisa
11 months agoLavonna
11 months agoGermaine
11 months agoCasie
11 months agoLenna
11 months agoGail
11 months agoBilly
11 months agoPatria
11 months agoMitsue
11 months agoSharen
11 months agoStaci
11 months ago