Mateus Oliveira dos Santos
Member since 2023
Diamond League
40229 points
Member since 2023
In this advanced challenge lab, you act as a Data Engineer for the Chicago Police Department. You will manage a high-stakes data integration project, moving IUCR reference data from Cloud Storage into BigQuery using code-free Dataproc Spark templates. Beyond simple ingestion, you will use BigQuery SQL to audit data quality, identify structural discrepancies like missing zero-padding, and reconcile mismatches between transaction records and reference tables to ensure analytical accuracy.
This lab tests your ability to develop a real-world Generative AI Q&A solution using a RAG framework. You will use Firestore as a vector database and deploy a Flask app as a user interface to query a food safety knowledge base.
In this advanced challenge lab, you act as a Data Engineer for Cymbal Direct, a retail company integrating real-time movie review data into a marketing pipeline. You are responsible for building two distinct streaming architectures. First, you will implement a direct, code-free ingestion path using Pub/Sub BigQuery subscriptions. Second, you will deploy a sophisticated Dataflow pipeline that uses JavaScript User-Defined Functions (UDFs) to transform raw text into numerical data before it reaches BigQuery, all while managing high-velocity data generated by a simulated stream.
In this course you will get hands-on in order to work through real-world challenges faced when building streaming data pipelines. The primary focus is on managing continuous, unbounded data with Google Cloud products.
This course demonstrates how to use AI/ML models for generative AI tasks in BigQuery. Through a practical use case involving customer relationship management, you learn the workflow of solving a business problem with Gemini models. To facilitate comprehension, the course also provides step-by-step guidance through coding solutions using both SQL queries and Python notebooks.
This course explores Gemini in BigQuery, a suite of AI-driven features to assist data-to-AI workflow. These features include data exploration and preparation, code generation and troubleshooting, and workflow discovery and visualization. Through conceptual explanations, a practical use case, and hands-on labs, the course empowers data practitioners to boost their productivity and expedite the development pipeline.
In this course, you learn about data engineering on Google Cloud, the roles and responsibilities of data engineers, and how those map to offerings provided by Google Cloud. You also learn about ways to address data engineering challenges.
Complete the intermediate Develop Serverless Applications on Cloud Run skill badge course to demonstrate skills in the following: integrating Cloud Run with Cloud Storage for data management, architecting resilient asynchronous systems using Cloud Run and Pub/Sub, constructing REST API gateways powered by Cloud Run, and building and deploying services on Cloud Run.
Complete the intermediate Manage Kubernetes in Google Cloud skill badge course to demonstrate skills in the following: managing deployments with kubectl, monitoring and debugging applications on Google Kubernetes Engine (GKE), and continuous delivery techniques.
Complete the intermediate Engineer AI Agents with Agent Development Kit (ADK) skill badge by completing this course to demonstrate skills in the following: formulating real-world language model research problems; building a simple tokenizer; preparing a dataset for training a transformer language model; running the training loop of a small language model.
Earn a skill badge by completing the Cloud Architecture: Design, Implement, and Manage to demonstrate skills in the following: deploy a publicly accessible website using Apache web servers, configure a Compute Engine VM using startup scripts, configure secure RDP using a Windows Bastion host and firewall rules, build and deploy a Docker image to a Kubernetes cluster and then update it, and create a CloudSQL instance and import a MySQL database. This skill badge is a great resource for understanding topics that will appear in the Google Cloud Certified Professional Cloud Architect certification exam.
Complete the intermediate Implement Cloud Security Fundamentals on Google Cloud skill badge course to demonstrate skills in the following: creating and assigning roles with Identity and Access Management (IAM); creating and managing service accounts; enabling private connectivity across virtual private cloud (VPC) networks; restricting application access using Identity-Aware Proxy; managing keys and encrypted data using Cloud Key Management Service (KMS); and creating a private Kubernetes cluster.
Google Cloud'da Uygulama Geliştirme Ortamı Oluşturma kursunu tamamlayarak beceri rozeti kazanın. Bu kursta Cloud Storage, Identity and Access Management, Cloud Functions ve Pub/Sub gibi teknolojilerin temel özelliklerini kullanarak depolama odaklı bulut altyapısı oluşturma ve bu altyapıyla bağlantı kurmayı öğreneceksiniz.
Üretken Yapay Zeka Ajanları: Kuruluşunuzu Dönüştürün, Üretken Yapay Zeka Lideri öğrenme rotasının beşinci ve son kursudur. Bu kursta, kuruluşların özel üretken yapay zeka ajanlarını kullanarak belirli işletme zorluklarının üstesinden nasıl gelebileceği ele alınmaktadır. Temel bir üretken yapay zeka ajanı oluşturarak pratik yapacak, bu ajanların modeller, mantık döngüleri ve araçlar gibi bileşenlerini keşfedeceksiniz.
Üretken Yapay Zeka Uygulamaları ile İşinizi Dönüştürün, Üretken Yapay Zeka Lideri öğrenme rotasının dördüncü kursudur. Bu kursta, Google'ın üretken yapay zeka uygulamaları (ör. Gemini ile Google Workspace ve NotebookLM) tanıtılmaktadır. Temellendirme, veriyle artırılmış üretim, etkili istemler hazırlama ve otomatik iş akışları oluşturma gibi kavramlar hakkında size rehberlik eder.
Üretken Yapay Zeka: Ekosistemi Tanıma, Üretken Yapay Zeka Lideri öğrenme rotasının üçüncü kursudur. Üretken yapay zeka, çalışma şeklimizi ve çevremizle etkileşim kurma biçimimizi değiştiriyor. Peki bir lider olarak bu teknolojinin gücünden yararlanıp işletmenizde nasıl gerçek sonuçlar elde edebilirsiniz? Bu kursta, üretken yapay zeka çözümleri oluşturmanın farklı katmanlarını, Google Cloud'un sunduğu hizmetleri ve çözüm seçerken dikkate alınması gereken faktörleri keşfedeceksiniz.
Üretken Yapay Zeka: Temel Kavramları Öğrenin, Üretken Yapay Zeka Lideri öğrenme rotasının ikinci kursudur. Bu kursta, yapay zeka, makine öğrenimi ve üretken yapay zeka arasındaki farkları keşfederek üretken yapay zekanın temel kavramlarını öğrenecek ve çeşitli veri türlerinin üretken yapay zekanın kurumsal zorlukları çözmesine nasıl yardımcı olduğunu anlayacaksınız. Temel modellerin sınırlamalarını gidermeye yardımcı olacak Google Cloud stratejileriyle sorumlu ve güvenli yapay zeka geliştirme ve dağıtımının temel zorlukları hakkında da bilgi edineceksiniz.
Üretken Yapay Zeka: Chatbot'tan Daha Fazlası, Üretken Yapay Zeka Lideri öğrenme rotasının ilk kursudur ve ön koşul gerektirmez. Bu kurs, chatbot'larla ilgili temel bilgilerin ötesine geçerek üretken yapay zekanın kuruluşunuza sağlayabileceği gerçek potansiyeli keşfetmeyi amaçlamaktadır. Üretken yapay zekanın gücünden yararlanmak için çok önemli olan temel modeller ve istem mühendisliği gibi kavramları keşfedeceksiniz. Kurs ayrıca kuruluşunuz için başarılı bir üretken yapay zeka stratejisi geliştirirken dikkate almanız gereken önemli noktalar hakkında size rehberlik edecek.
Complete the introductory Create and Manage Cloud SQL for PostgreSQL Instances skill badge to demonstrate skills in the following: migrating, configuring, and managing Cloud SQL for PostgreSQL instances and databases.
Complete the introductory Create and Manage Cloud Spanner Instances skill badge to demonstrate skills in the following: creating and interacting with Cloud Spanner instances and databases; loading Cloud Spanner databases using various techniques; backing up Cloud Spanner databases; defining schemas and understanding query plans; and deploying a Modern Web App connected to a Cloud Spanner instance.
Complete the introductory Create and Manage AlloyDB Instances skill badge to demonstrate skills in the following: performing core AlloyDB operations and tasks, migrating to AlloyDB from PostgreSQL, administering an AlloyDB database, and accelerating analytical queries using the AlloyDB Columnar Engine.
Complete the intermediate Engineer Data for Predictive Modeling with BigQuery ML skill badge to demonstrate skills in the following: building data transformation pipelines to BigQuery using Dataprep by Trifacta; using Cloud Storage, Dataflow, and BigQuery to build extract, transform, and load (ETL) workflows; and building machine learning models using BigQuery ML.
In the last installment of the Dataflow course series, we will introduce the components of the Dataflow operational model. We will examine tools and techniques for troubleshooting and optimizing pipeline performance. We will then review testing, deployment, and reliability best practices for Dataflow pipelines. We will conclude with a review of Templates, which makes it easy to scale Dataflow pipelines to organizations with hundreds of users. These lessons will help ensure that your data platform is stable and resilient to unanticipated circumstances.
Complete the introductory Build a Data Mesh with Knowledge Catalog skill badge to demonstrate skills in the following: building a data mesh with Knowledge Catalog to facilitate data security, governance, and discovery on Google Cloud. You practice and test your skills in tagging assets, assigning IAM roles, and assessing data quality in Knowledge Catalog.
Complete the intermediate Build a Data Warehouse with BigQuery skill badge course to demonstrate skills in the following: joining data to create new tables, troubleshooting joins, appending data with unions, creating date-partitioned tables, and working with JSON, arrays, and structs in BigQuery.
In this second installment of the Dataflow course series, we are going to be diving deeper on developing pipelines using the Beam SDK. We start with a review of Apache Beam concepts. Next, we discuss processing streaming data using windows, watermarks and triggers. We then cover options for sources and sinks in your pipelines, schemas to express your structured data, and how to do stateful transformations using State and Timer APIs. We move onto reviewing best practices that help maximize your pipeline performance. Towards the end of the course, we introduce SQL and Dataframes to represent your business logic in Beam and how to iteratively develop pipelines using Beam notebooks.
Giriş düzeyindeki Compute Engine İçin Cloud Load Balancing'i Uygulama beceri rozetini tamamlayarak şu konulardaki becerilerinizi gösterin: Compute Engine'de sanal makineler oluşturma ve dağıtma. Ağ ve uygulama yük dengeleyicileri yapılandırma.
This course is part 1 of a 3-course series on Serverless Data Processing with Dataflow. In this first course, we start with a refresher of what Apache Beam is and its relationship with Dataflow. Next, we talk about the Apache Beam vision and the benefits of the Beam Portability framework. The Beam Portability framework achieves the vision that a developer can use their favorite programming language with their preferred execution backend. We then show you how Dataflow allows you to separate compute and storage while saving money, and how identity, access, and management tools interact with your Dataflow pipelines. Lastly, we look at how to implement the right security model for your use case on Dataflow.
Incorporating machine learning into data pipelines increases the ability to extract insights from data. This course covers ways machine learning can be included in data pipelines on Google Cloud. For little to no customization, this course covers AutoML. For more tailored machine learning capabilities, this course introduces Notebooks and BigQuery machine learning (BigQuery ML). Also, this course covers how to productionalize machine learning solutions by using Vertex AI.
This 1-week, accelerated on-demand course builds upon Google Cloud Platform Big Data and Machine Learning Fundamentals. Through a combination of video lectures, demonstrations, and hands-on labs, you'll learn to build streaming data pipelines using Google cloud Pub/Sub and Dataflow to enable real-time decision making. You will also learn how to build dashboards to render tailored output for various stakeholder audiences.
In this intermediate course, you will learn to design, build, and optimize robust batch data pipelines on Google Cloud. Moving beyond fundamental data handling, you will explore large-scale data transformations and efficient workflow orchestration, essential for timely business intelligence and critical reporting. Get hands-on practice using Dataflow for Apache Beam and Serverless for Apache Spark (Dataproc Serverless) for implementation, and tackle crucial considerations for data quality, monitoring, and alerting to ensure pipeline reliability and operational excellence. A basic knowledge of data warehousing, ETL/ELT, SQL, Python, and Google Cloud concepts is recommended.
Giriş düzeyindeki Compute Engine İçin Cloud Load Balancing'i Uygulama beceri rozetini tamamlayarak şu konulardaki becerilerinizi gösterin: Compute Engine'de sanal makineler oluşturma ve dağıtma. Ağ ve uygulama yük dengeleyicileri yapılandırma.
Giriş düzeyindeki Google Cloud'da Makine Öğrenimi API'leri İçin Veri Hazırlama beceri rozetini tamamlayarak şu konulardaki becerilerinizi gösterin: Dataprep by Trifacta ile veri temizleme, Dataflow'da veri ardışık düzenleri çalıştırma, Managed Service for Apache Spark'ta küme oluşturma ve Apache Spark işleri çalıştırma ve makine öğrenimi API'lerini (Cloud Natural Language API, Google Cloud Speech-to-Text API ve Video Intelligence API dahil olmak üzere) çağırma.
While the traditional approaches of using data lakes and data warehouses can be effective, they have shortcomings, particularly in large enterprise environments. This course introduces the concept of a data lakehouse and the Google Cloud products used to create one. A lakehouse architecture uses open-standard data sources and combines the best features of data lakes and data warehouses, which addresses many of their shortcomings.
This course introduces the Google Cloud big data and machine learning products and services that support the data-to-AI lifecycle. It explores the processes, challenges, and benefits of building a big data pipeline and machine learning models with Vertex AI on Google Cloud.
This course helps learners create a study plan for the PDE (Professional Data Engineer) certification exam. Learners explore the breadth and scope of the domains covered in the exam. Learners assess their exam readiness and create their individual study plan.