Deal of The Day! Hurry Up, Grab the Special Discount - Save 25% - Ends In 00:00:00 Coupon code: SAVE25
Welcome to Pass4Success

- Free Preparation Discussions

Amazon AIF-C01 Exam - Topic 4 Question 40 Discussion

An education company wants to build a private tutor application. The application will give users the ability to enter text or provide a picture of a question. The application will respond with a written answer and an explanation of the written answer.Which model type meets these requirements?
B) Multimodal LLM
A) Computer vision model
C) Diffusion model
D) Text-to-speech model

Amazon AIF-C01 Exam - Topic 4 Question 40 Discussion

Actual exam question for Amazon's AIF-C01 exam
Question #: 40
Topic #: 4
[All AIF-C01 Questions]

An education company wants to build a private tutor application. The application will give users the ability to enter text or provide a picture of a question. The application will respond with a written answer and an explanation of the written answer.

Which model type meets these requirements?

Show Suggested Answer Hide Answer
Suggested Answer: B

Comprehensive and Detailed Explanation From Exact AWS AI documents:

A multimodal large language model (LLM) can:

Accept both text and image inputs

Understand visual and textual context

Generate coherent written explanations

AWS generative AI guidance positions multimodal LLMs as the best choice for applications requiring cross-modal understanding and text generation.

Why the other options are incorrect:

Computer vision (A) does not generate text explanations.

Diffusion models (C) generate images.

Text-to-speech (D) converts text to audio.

AWS AI document references:

Multimodal Foundation Models on AWS

Building AI Tutors with Generative Models


Contribute your Thoughts:

0/2000 characters
Jesusa
1 day ago
I agree with Candra, B is the best fit for this app.
upvoted 0 times
...
Sage
7 days ago
Wait, can it really analyze images and text together? Sounds tricky.
upvoted 0 times
...
Avery
12 days ago
No way, it has to be B) Multimodal LLM!
upvoted 0 times
...
Yaeko
17 days ago
I think A) Computer vision model could work too, but not as well.
upvoted 0 times
...
Candra
22 days ago
Definitely B) Multimodal LLM, it handles both text and images.
upvoted 0 times
...
Miss
27 days ago
I’m leaning towards the multimodal LLM as well, but I’m not confident. I just hope I remember the differences clearly during the exam!
upvoted 0 times
...
Marget
1 month ago
I feel like the multimodal LLM is the best fit here, but I wonder if there's a chance the diffusion model could work too?
upvoted 0 times
...
Sharen
1 month ago
I remember practicing a question about model types, and I think a computer vision model might only handle images, so that doesn't seem right.
upvoted 0 times
...
Wendell
1 month ago
I think the multimodal LLM could be the right choice since it can handle both text and images, but I'm not entirely sure.
upvoted 0 times
...

Save Cancel