Blog - Albert Nogués

Data Quality Checks with Soda-Core in Databricks

updated on May 31, 2024May 31, 2024

It’s easy to do data quality checks when working with spark with the soda-core library. The library has support for spark dataframes. I’ve tested it within a databricks environment and it worked quite easily for me. For the examples of this article i am loading the customers table from the tpch delta tables in the …

Query Delta Tables in the DataLake from PowerBi with Databricks

updated on November 16, 2023November 15, 2023

There are several ways to query delta tables from PowerBi. We are going to cover the 4th method here. To do it first we need a service princpal, a secret scope pointing to a databricks keyvault and the password of the SPN stored in this keyvault. Once we have this, the first step is to …

Databricks query federation with Snowflake. Easy and Fast!

updated on January 31, 2023January 31, 2023

Introduction In the same way that is possible to read and write data from snowflake inside databricks, its also possible to use databricks with query federation against diverse SQL engines, including snowflake. The current supported engines are: We are going to demonstrate how it works with Snowflake. We will first create a table in databricks, …

Useful Databricks/Spark resources

updated on December 14, 2022December 14, 2022

Memory Profiling in PySpark: https://www.databricks.com/blog/2022/11/30/memory-profiling-pyspark.html Run Databricks queries directly from VSCODE: https://ganeshchandrasekaran.com/run-your-databricks-sql-queries-from-vscode-9c70c5d4903c Spark Testing with chispa: https://github.com/alexott/spark-playground/tree/master/testing Best Practices for Cost Management on Databricks: https://www.databricks.com/blog/2022/10/18/best-practices-cost-management-databricks.html UDF Pyspark: https://docs.databricks.com/udf/python.html Pandas UDF’s: https://docs.databricks.com/udf/pandas.html Introducing Pandas UDF for PySpark: https://www.databricks.com/blog/2017/10/30/introducing-vectorized-udfs-for-pyspark.html

Smallest Analytical Platform Ever!

updated on May 7, 2022May 7, 2022

I’ve started working on some of my free time in a project to build the smallest useful analytics platform on the cloud (starting with azure). The purpose is to use it a sa PoC to show to colleagues, managers, prospective customers or just to have fun and play It’s publicly available on my github repo …

Implementing CI/CD in Databricks with Azure DevOps (Part 1)

updated on April 30, 2022April 30, 2022

There are many ways to implement some CI/CD with Databricks. We can use Azure DevOps, Github+Github Actions or any other combination of tools, including the dbx tool. But an easy way to just copy notebooks between workspaces can be implemented easily with Azure DevOps. We are going to use the git repos capability of Azure …

Centos 8/9 Stream and AlmaLinux images for WSL

updated on April 12, 2022April 12, 2022

For the ones like me, interested in running Linux systems on windows for many automation or administration tasks, I am sharing here the images i’ve found: Centos7/8/9 Stream: https://github.com/mishamosher/CentOS-WSL AlmaLinux (Centos Replacement equivalent to RHEL): https://www.microsoft.com/en-us/p/almalinux-8-wsl/9nmd96xjj19f#activetab=pivot:overviewtab The latest one is a direct link to the microsoft store.

Using Azure Private Endpoints with Databricks

updated on November 22, 2023December 9, 2021

In this article i will show how to avoing going outside to the internet when using resources inside azure, specially if they are in the same subscription and location (datacenter). Why we may want a private endpoint? Thats a good question. For oth security and performance. Just like using TSCM Equipment for optimal safety and …

Databricks connectivity to Azure SQL / SQL Server

updated on December 9, 2021December 9, 2021

Most of the developments I see inside databricks rely on fetching or writing data to some sort of Database. Usually the preferred method for this is though the use of jdbc driver, as most databases offer some sort of jdbc driver. In some cases, though, its also possible to use some spark optimized driver. This …