Running Spark on Kubernetes with Dataproc

https://storage.googleapis.com/gweb-cloudblog-publish/images/19_-_Infrastructure_Modernization_o5CKMmf.max-2600x2600.jpg

Apache Spark is now the de facto standard for data engineering, data exploration and machine learning. Just likeKubernetes (k8s), is for automating containerized application deployment, scaling, and management. The open source ecosystem is now converging towards utilizing k8s as the compute platform in addition to YARN.

Today, we are announcing the general availability of Dataproc on Google Kubernetes Engine (GKE), enabling you to leverage k8s to manage and optimize your compute platforms. You can now create a Dataproc cluster and submit Spark jobs on a self-managed GKE cluster.

Dataproc on GKE for Spark (GA)

K8s builds on 15 years of running Google's containerized workloads and the critical contributions from the open source community. Inspired by Google’s internal cluster management system, Borg, K8s makes everything associated with deploying and managing your application easier. With the widespread adoption of k8s, many customers are now standardizing on k8s for their compute platform...

Copyright of this story solely belongs to cloud.google.com. To see the full text click HERE

Read more