Deal of The Day! Hurry Up, Grab the Special Discount - Save 25% - Ends In 00:00:00 Coupon code: SAVE25
Welcome to Pass4Success

- Free Preparation Discussions

Databricks Certified Data Engineer Associate Exam - Topic 1 Question 69 Discussion

A data engineering team needs to integrate two data sources into Databricks:Clickstream events: 5,000 events per second from an Apache Kafka topicCustomer master data: Only changed records every four hours from a Snowflake databaseThe solution must process clickstream data with latency under 30 seconds and prevent reprocessing customer master data that has not changed.Which ingestion approach meets these requirements?
A) Use Structured Streaming for Kafka and a Lakeflow Connect managed connector with incremental processing for Snowflake.
B) Use a non-incremental Snowflake connector, fetch all data every four hours, and apply MERGE operations.
C) Use spark.readStream() with Kafka and query Snowflake hourly using a time-based filter.
D) Use Structured Streaming for Kafka and a Lakeflow Connect connector with a full refresh for Snowflake.

Databricks Certified Data Engineer Associate Exam - Topic 1 Question 69 Discussion

Actual exam question for Databricks's Databricks Certified Data Engineer Associate exam
Question #: 69
Topic #: 1
[All Databricks Certified Data Engineer Associate Questions]

A data engineering team needs to integrate two data sources into Databricks:

Clickstream events: 5,000 events per second from an Apache Kafka topic

Customer master data: Only changed records every four hours from a Snowflake database

The solution must process clickstream data with latency under 30 seconds and prevent reprocessing customer master data that has not changed.

Which ingestion approach meets these requirements?

Show Suggested Answer Hide Answer
Suggested Answer: A

Contribute your Thoughts:

0/2000 characters
Flo
5 days ago
C) could work, but not sure about the hourly query part.
upvoted 0 times
...
Glen
10 days ago
B) seems inefficient, fetching all data every time? No thanks.
upvoted 0 times
...
Aleisha
15 days ago
Wait, can we really avoid reprocessing with A)? Sounds risky.
upvoted 0 times
...
Han
20 days ago
Totally agree, A) makes the most sense here!
upvoted 0 times
...
Alonso
25 days ago
A) is the best option for low latency with Kafka.
upvoted 0 times
...
Tarra
1 month ago
I think option D is not ideal because a full refresh for Snowflake could lead to unnecessary data processing, which we wanted to avoid.
upvoted 0 times
...
Brittni
1 month ago
Option C sounds familiar, but I’m concerned about the latency requirement. I feel like querying Snowflake hourly might not meet the 30-second limit.
upvoted 0 times
...
Christiane
1 month ago
I'm not entirely sure about the Snowflake connector; I think we practiced a similar question where we had to avoid fetching all data unnecessarily.
upvoted 0 times
...
Casie
2 months ago
I remember we discussed using Structured Streaming for real-time data, so option A seems like a good fit for the Kafka part.
upvoted 0 times
...

Save Cancel