Deal of The Day! Hurry Up, Grab the Special Discount - Save 25% - Ends In 00:00:00 Coupon code: SAVE25
Welcome to Pass4Success

- Free Preparation Discussions

Databricks Certified Associate Developer for Apache Spark 3.5 Exam - Topic 4 Question 21 Discussion

46 of 55.A data engineer is implementing a streaming pipeline with watermarking to handle late-arriving records.The engineer has written the following code:inputStream \.withWatermark("event_time", "10 minutes") \.groupBy(window("event_time", "15 minutes"))What happens to data that arrives after the watermark threshold?
A) Any data arriving more than 10 minutes after the watermark threshold will be ignored and not included in the aggregation.
B) Records that arrive later than the watermark threshold (10 minutes) will automatically be included in the aggregation if they fall within the 15-minute window.
C) Data arriving more than 10 minutes after the latest watermark will still be included in the aggregation but will be placed into the next window.
D) The watermark ensures that late data arriving within 10 minutes of the latest event time will be processed and included in the windowed aggregation.

Databricks Certified Associate Developer for Apache Spark 3.5 Exam - Topic 4 Question 21 Discussion

Actual exam question for Databricks's Databricks Certified Associate Developer for Apache Spark 3.5 exam
Question #: 21
Topic #: 4
[All Databricks Certified Associate Developer for Apache Spark 3.5 Questions]

46 of 55.

A data engineer is implementing a streaming pipeline with watermarking to handle late-arriving records.

The engineer has written the following code:

inputStream \

.withWatermark("event_time", "10 minutes") \

.groupBy(window("event_time", "15 minutes"))

What happens to data that arrives after the watermark threshold?

Show Suggested Answer Hide Answer
Suggested Answer: A

Watermarking in Structured Streaming defines how late a record can arrive based on event time before Spark discards it.

Behavior:

.withWatermark('event_time', '10 minutes')

This means Spark will keep state for 10 minutes beyond the maximum event time seen so far.

Any data arriving later than 10 minutes after the current watermark is ignored --- it will not be included in the aggregation or output.

Why the other options are incorrect:

B: Late data beyond the watermark threshold is not included.

C: Late data is not moved to a new window; it's simply dropped.

D: True for late data within the watermark threshold, not after it.


Spark Structured Streaming Guide --- withWatermark() behavior and late data handling.

Databricks Exam Guide (June 2025): Section ''Structured Streaming'' --- watermarking and state cleanup behavior.

===========

Contribute your Thoughts:

0/2000 characters

Currently there are no comments in this discussion, be the first to comment!


Save Cancel