Google Professional Machine Learning Engineer Exam - Topic 9 Question 23 Discussion
You developed an ML model with Al Platform, and you want to move it to production. You serve a few thousand queries per second and are experiencing latency issues. Incoming requests are served by a load balancer that distributes them across multiple Kubeflow CPU-only pods running on Google Kubernetes Engine (GKE). Your goal is to improve the serving latency without changing the underlying infrastructure. What should you do?
D) Recompile TensorFlow Serving using the source to support CPU-specific optimizations Instruct GKE to choose an appropriate baseline minimum CPU platform for serving nodes
A) Significantly increase the max_batch_size TensorFlow Serving parameter
B) Switch to the tensorflow-model-server-universal version of TensorFlow Serving
C) Significantly increase the max_enqueued_batches TensorFlow Serving parameter
Rusty
10 months agoLashawnda
10 months agoRoosevelt
11 months agoCherelle
11 months agoLoreen
11 months agoKimberlie
11 months agoJutta
11 months agoViola
11 months agoGalen
11 months agoLoreta
11 months agoCorrinne
11 months agoLeanora
11 months agoCarissa
11 months ago