Databricks Certified Data Engineer Associate Exam - Topic 1 Question 69 Discussion
A data engineering team needs to integrate two data sources into Databricks:Clickstream events: 5,000 events per second from an Apache Kafka topicCustomer master data: Only changed records every four hours from a Snowflake databaseThe solution must process clickstream data with latency under 30 seconds and prevent reprocessing customer master data that has not changed.Which ingestion approach meets these requirements?
A) Use Structured Streaming for Kafka and a Lakeflow Connect managed connector with incremental processing for Snowflake.
B) Use a non-incremental Snowflake connector, fetch all data every four hours, and apply MERGE operations.
C) Use spark.readStream() with Kafka and query Snowflake hourly using a time-based filter.
D) Use Structured Streaming for Kafka and a Lakeflow Connect connector with a full refresh for Snowflake.
Flo
5 days agoGlen
10 days agoAleisha
15 days agoHan
20 days agoAlonso
25 days agoTarra
1 month agoBrittni
1 month agoChristiane
1 month agoCasie
2 months ago