<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Spark Archives - Albert Nogués</title>
	<atom:link href="https://www.albertnogues.com/category/spark/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.albertnogues.com/category/spark/</link>
	<description>Data and Cloud Freelancer</description>
	<lastBuildDate>Wed, 22 Nov 2023 10:18:52 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	

<image>
	<url>https://www.albertnogues.com/wp-content/uploads/2020/12/cropped-cropped-AlbertLogo2-32x32.png</url>
	<title>Spark Archives - Albert Nogués</title>
	<link>https://www.albertnogues.com/category/spark/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Useful Databricks/Spark resources</title>
		<link>https://www.albertnogues.com/useful-databricks-spark-resources/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=useful-databricks-spark-resources</link>
		
		<dc:creator><![CDATA[Albert]]></dc:creator>
		<pubDate>Wed, 14 Dec 2022 12:58:28 +0000</pubDate>
				<category><![CDATA[BigData]]></category>
		<category><![CDATA[Databricks]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Spark]]></category>
		<category><![CDATA[SQL]]></category>
		<guid isPermaLink="false">https://www.albertnogues.com/?p=1694</guid>

					<description><![CDATA[<p>Memory Profiling in PySpark: https://www.databricks.com/blog/2022/11/30/memory-profiling-pyspark.html Run Databricks queries directly from VSCODE: https://ganeshchandrasekaran.com/run-your-databricks-sql-queries-from-vscode-9c70c5d4903c Spark Testing with chispa: https://github.com/alexott/spark-playground/tree/master/testing Best Practices for Cost Management on Databricks: https://www.databricks.com/blog/2022/10/18/best-practices-cost-management-databricks.html UDF Pyspark: https://docs.databricks.com/udf/python.html Pandas UDF&#8217;s: https://docs.databricks.com/udf/pandas.html Introducing Pandas UDF for PySpark: https://www.databricks.com/blog/2017/10/30/introducing-vectorized-udfs-for-pyspark.html</p>
<p>The post <a href="https://www.albertnogues.com/useful-databricks-spark-resources/">Useful Databricks/Spark resources</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p>Memory Profiling in PySpark: <a href="https://www.databricks.com/blog/2022/11/30/memory-profiling-pyspark.html" target="_blank" rel="noreferrer noopener">https://www.databricks.com/blog/2022/11/30/memory-profiling-pyspark.html</a></p>



<p>Run Databricks queries directly from VSCODE: <a href="https://ganeshchandrasekaran.com/run-your-databricks-sql-queries-from-vscode-9c70c5d4903c" target="_blank" rel="noreferrer noopener">https://ganeshchandrasekaran.com/run-your-databricks-sql-queries-from-vscode-9c70c5d4903c</a></p>



<p>Spark Testing with chispa: <a href="https://github.com/alexott/spark-playground/tree/master/testing" target="_blank" rel="noreferrer noopener">https://github.com/alexott/spark-playground/tree/master/testing</a></p>



<p>Best Practices for Cost Management on Databricks: <a href="https://www.databricks.com/blog/2022/10/18/best-practices-cost-management-databricks.html" target="_blank" rel="noreferrer noopener">https://www.databricks.com/blog/2022/10/18/best-practices-cost-management-databricks.html</a></p>



<p>UDF Pyspark: <a href="https://docs.databricks.com/udf/python.html" target="_blank" rel="noreferrer noopener">https://docs.databricks.com/udf/python.html</a></p>



<p>Pandas UDF&#8217;s: <a href="https://docs.databricks.com/udf/pandas.html" target="_blank" rel="noreferrer noopener">https://docs.databricks.com/udf/pandas.html</a></p>



<p>Introducing Pandas UDF for PySpark: <a href="https://www.databricks.com/blog/2017/10/30/introducing-vectorized-udfs-for-pyspark.html" target="_blank" rel="noreferrer noopener">https://www.databricks.com/blog/2017/10/30/introducing-vectorized-udfs-for-pyspark.html</a></p>
<p>The post <a href="https://www.albertnogues.com/useful-databricks-spark-resources/">Useful Databricks/Spark resources</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Smallest Analytical Platform Ever!</title>
		<link>https://www.albertnogues.com/smallest-analytical-platform-ever/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=smallest-analytical-platform-ever</link>
		
		<dc:creator><![CDATA[Albert]]></dc:creator>
		<pubDate>Sat, 07 May 2022 08:38:12 +0000</pubDate>
				<category><![CDATA[Azure]]></category>
		<category><![CDATA[Cloud]]></category>
		<category><![CDATA[Databricks]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Spark]]></category>
		<category><![CDATA[SQL]]></category>
		<category><![CDATA[databricks]]></category>
		<category><![CDATA[git]]></category>
		<category><![CDATA[python]]></category>
		<category><![CDATA[spark]]></category>
		<guid isPermaLink="false">https://www.albertnogues.com/?p=1440</guid>

					<description><![CDATA[<p>I&#8217;ve started working on some of my free time in a project to build the smallest useful analytics platform on the cloud (starting with azure). The purpose is to use it a sa PoC to show to colleagues, managers, prospective customers or just to have fun and play It&#8217;s publicly available on my github repo &#8230; </p>
<p>The post <a href="https://www.albertnogues.com/smallest-analytical-platform-ever/">Smallest Analytical Platform Ever!</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p>I&#8217;ve started working on some of my free time in a project to build the smallest useful analytics platform on the cloud (starting with azure).</p>



<p>The purpose is to use it a sa PoC to show to colleagues, managers, prospective customers or just to have fun and play</p>



<p>It&#8217;s publicly available on my github repo and any collaboration is welcome. You can fork it, improve it, send PR&#8217;s and do whatever you want!</p>



<p>The first version will run solely on azure. The objective is to show the following technologies/disciplines:</p>



<p>* Infrastructure as a Code (IaaC), by using Terraform</p>



<p>* Cloud architecture anc Cloud Ops by using an azure cloud environment</p>



<p>* Data Engineering by using a Spark powered Databricks Notebook and an ADF Pipeline (future)</p>



<p>* DevOps to trigger some pipelines based on changes (future)</p>



<p>* Basic Security concepts (keyvault, service principals, least privileged rbac accesses&#8230;)</p>



<p>* FinOps keeping the costs at minimum and choosing the proper tools for the job</p>



<p>* Reporting and Dashboarding on data in the platform</p>



<p>* Data management: We will use an adls storage account and azure sql db</p>



<p>TOOLS:</p>



<p>* Terraform to deploy all the infra as a code</p>



<p>* Azure Cloud to host our resources</p>



<p>You have the code plus all the information on my github repo:</p>



<p><a href="https://github.com/anogues/ProjectZ">https://github.com/anogues/ProjectZ</a></p>



<p></p>
<p>The post <a href="https://www.albertnogues.com/smallest-analytical-platform-ever/">Smallest Analytical Platform Ever!</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Implementing CI/CD in Databricks with Azure DevOps (Part 1)</title>
		<link>https://www.albertnogues.com/implementing-ci-cd-in-databricks-with-azure-devops-part-1/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=implementing-ci-cd-in-databricks-with-azure-devops-part-1</link>
		
		<dc:creator><![CDATA[Albert]]></dc:creator>
		<pubDate>Sat, 30 Apr 2022 15:10:29 +0000</pubDate>
				<category><![CDATA[Azure]]></category>
		<category><![CDATA[Databricks]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Spark]]></category>
		<category><![CDATA[Cloud]]></category>
		<category><![CDATA[databricks]]></category>
		<category><![CDATA[git]]></category>
		<category><![CDATA[spark]]></category>
		<guid isPermaLink="false">https://www.albertnogues.com/?p=1416</guid>

					<description><![CDATA[<p>There are many ways to implement some CI/CD with Databricks. We can use Azure DevOps, Github+Github Actions or any other combination of tools, including the dbx tool. But an easy way to just copy notebooks between workspaces can be implemented easily with Azure DevOps. We are going to use the git repos capability of Azure &#8230; </p>
<p>The post <a href="https://www.albertnogues.com/implementing-ci-cd-in-databricks-with-azure-devops-part-1/">Implementing CI/CD in Databricks with Azure DevOps (Part 1)</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p>There are many ways to implement some CI/CD with Databricks. We can use Azure DevOps, Github+Github Actions or any other combination of tools, including the <a href="https://dbx.readthedocs.io/en/latest/templates/python_basic.html#project-file-structure" target="_blank" rel="noreferrer noopener">dbx tool</a>.</p>



<p>But an easy way to just copy notebooks between workspaces can be implemented easily with Azure DevOps.</p>



<p>We are going to use the git repos capability of Azure Databricks, so when a new code change is commited in a notebook an Azure DevOps pipeline will trigger the transport copy of the workbook from the first Databricks workspace (in our case a NonProd workspace) to the target one, again, in our case, the Prod workspace.</p>



<p>To achieve this we will use some more components from the Azure ecosystem, including the use of Keyvaults to keep all our secrets stored safely. The list of prerequisites is the following:</p>



<ul class="wp-block-list"><li>Two Databricks workspaces, one our source workspace (NonProd) and another, our Production one.</li><li>An Azure Keyvault (or two if we want to segregate the environments)</li><li>Azure Databricks repository configured at least in our source workspace, so when the change is commited we can triger the pipeline that will fetch the notebook and transport it to the prod workspace</li><li>Access to Azure DevOps (Something similar can be implemented with Github + Github Actions)</li></ul>



<p>Lets see how to implement it. First we need to make sure git repos in enabled our source workspace. We can verify by login with an admin privileged user to our workspace and make sure the option is checked as follows:</p>



<figure class="wp-block-image size-large is-resized"><img fetchpriority="high" decoding="async" src="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD1-924x1024.png" alt="" class="wp-image-1418" width="687" height="760" srcset="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD1-924x1024.png 924w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD1-271x300.png 271w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD1-768x851.png 768w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD1-54x60.png 54w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD1.png 1072w" sizes="(max-width: 687px) 100vw, 687px" /><figcaption>Fig 1. Make sure that github repos is enabled in our workspace.</figcaption></figure>



<p>Secondly, we go to <a href="https://azure.microsoft.com/en-us/services/devops/" target="_blank" rel="noreferrer noopener">Azure DevOps services</a> and we create a new project. I&#8217;ve called it DatabricksCICD but feel free to call it whatever you need:</p>



<figure class="wp-block-image size-large"><img decoding="async" width="1024" height="714" src="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD2-1024x714.png" alt="" class="wp-image-1419" srcset="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD2-1024x714.png 1024w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD2-300x209.png 300w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD2-768x536.png 768w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD2-86x60.png 86w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD2.png 1531w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption>Fig 2. Create a new repo and initialize it</figcaption></figure>



<p>Once created we take the details for cloning our repo and copying them. We go to our databricks workspace and then we look for the Repos option on the left, and add a new repository. We need to paste the url to clone our newle Azure DevOps created repository:</p>



<figure class="wp-block-image size-large"><img decoding="async" width="1024" height="319" src="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD3-1024x319.png" alt="" class="wp-image-1420" srcset="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD3-1024x319.png 1024w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD3-300x93.png 300w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD3-768x239.png 768w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD3-1536x478.png 1536w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD3-2048x637.png 2048w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD3-193x60.png 193w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption>Fig 3. Cloning our Azure DevOps Repository</figcaption></figure>



<p>Once we linked our Databricks workspace with our DevOps repo, now we can create a new notebook. In the same Repos section, click on the down arrow to create a new notebook, as shown below:</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="889" height="373" src="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD4.png" alt="" class="wp-image-1421" srcset="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD4.png 889w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD4-300x126.png 300w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD4-768x322.png 768w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD4-143x60.png 143w" sizes="auto, (max-width: 889px) 100vw, 889px" /><figcaption>Fig 4.  Creating a new notebook.</figcaption></figure>



<p>The content of the notebook, you can put anything you want. I&#8217;m writing a print(&#8220;Hello from Albert&#8221;) statement. We will not run it, we just want to show it&#8217;s possible to transport it. Once done, click on the save now in the revision tab:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="73" src="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD5-1024x73.png" alt="" class="wp-image-1423" srcset="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD5-1024x73.png 1024w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD5-300x21.png 300w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD5-768x55.png 768w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD5-1536x110.png 1536w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD5-2048x146.png 2048w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD5-600x43.png 600w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption>Fig 5. Saving our changes to the notebook.</figcaption></figure>



<p>Then click on the left in the main branch button, from there we will be commiting the changes to our repository:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="314" src="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD6-1024x314.png" alt="" class="wp-image-1424" srcset="https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD6-1024x314.png 1024w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD6-300x92.png 300w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD6-768x235.png 768w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD6-1536x470.png 1536w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD6-2048x627.png 2048w, https://www.albertnogues.com/wp-content/uploads/2022/04/DatabricksCICD6-196x60.png 196w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption>Fig 6. Pushing our notebook to the DevOps repo</figcaption></figure>



<p>If we go back now to our Azure DevOps project we should see the file has been commited to the repository. This ends the first part of this tutorial.</p>



<p>In the second blog entry we will see how to trigger the pipeline after a modification of this notebook and passing the credentials of the second workspace to be able to deliver the changed notebook to our Production (target) workspace.</p>
<p>The post <a href="https://www.albertnogues.com/implementing-ci-cd-in-databricks-with-azure-devops-part-1/">Implementing CI/CD in Databricks with Azure DevOps (Part 1)</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Using Azure Private Endpoints with Databricks</title>
		<link>https://www.albertnogues.com/using-azure-private-endpoints-with-databricks/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=using-azure-private-endpoints-with-databricks</link>
		
		<dc:creator><![CDATA[Albert]]></dc:creator>
		<pubDate>Thu, 09 Dec 2021 19:31:32 +0000</pubDate>
				<category><![CDATA[Azure]]></category>
		<category><![CDATA[Databricks]]></category>
		<category><![CDATA[Spark]]></category>
		<category><![CDATA[SQL]]></category>
		<category><![CDATA[Cloud]]></category>
		<category><![CDATA[databricks]]></category>
		<category><![CDATA[PrivateEndpoints]]></category>
		<category><![CDATA[spark]]></category>
		<guid isPermaLink="false">https://www.albertnogues.com/?p=1235</guid>

					<description><![CDATA[<p>In this article i will show how to avoing going outside to the internet when using resources inside azure, specially if they are in the same subscription and location (datacenter). Why we may want a private endpoint? Thats a good question. For oth security and performance. Just like using TSCM Equipment for optimal safety and &#8230; </p>
<p>The post <a href="https://www.albertnogues.com/using-azure-private-endpoints-with-databricks/">Using Azure Private Endpoints with Databricks</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p>In this article i will show how to avoing going outside to the internet when using resources inside azure, specially if they are in the same subscription and location (datacenter).</p>



<p>Why we may want a private endpoint? Thats a good question. For oth security and performance. Just like using <a href="https://spyassociates.com/counter-surveillance">TSCM Equipment</a> for optimal safety and security. We dont want the traffic going outside to the internet to return again back to the azure datacenter if the resource we are trying to reach is already there. So with a PrivateLink the traffic will stay inside the Azure backbone network avoiding reaching the internet. More information about private endpoints <a href="https://azure.microsoft.com/en-us/services/private-link/" target="_blank" rel="noreferrer noopener">here</a> and <a href="https://docs.microsoft.com/en-us/azure/private-link/private-link-overview" target="_blank" rel="noreferrer noopener">here</a>.</p>



<p>Though its possible to create private endpoints to connect to services in other subcriptions we will use the same subscription and the West Europe Region in this article. The goal is to connect to both a AzureSQL database using private connectivity and to a datalake using private connectivity as well.</p>



<h2 class="wp-block-heading">Creating a Private Endpoint for AzureSQL and integrating in the databricks vnet</h2>



<p>For this, i created a Databricks workspace and selected to use an already existing VNET, so this way I can add a new subnet for my private endpoints. One of the good things of doing this way is that NICs between subnets see each other and are reacheable (unless we block it with a network security group) but by default traffic is open within the VNET. So I can create a Private endpoint in a specific subnet of the same VNET that hosts the databricks subnets.</p>



<p>Bear in mind that it&#8217;s not possible to add a private endpoint to a subnet managed by databricks. So the two subnets we created, when we deployed the databricks workspace (Bot public and private) should not be modified. We will create a new one as shown in the screen below:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="145" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep4-1024x145.png" alt="" class="wp-image-1236" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep4-1024x145.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep4-300x43.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep4-768x109.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep4-1536x218.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep4-2048x290.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep4-424x60.png 424w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Our  Databricks VNET. Among the two subnets created when the databricks workspace is created i added a new one to host our Private Endpoints</figcaption></figure>



<p>Once defined properly the VNET, we are going to create the private endpoint to reach our AzureSQL through it.</p>



<p>First we need to go to the Azure Portal, find our AzureSQL Server, and click on the left menu called Private Endpoint Connections and click on the plus sing on top to create a new one. We just need to select the subscription, the resource group the nameof the private endpoint and the region. We can fill it as shown in the following picture:</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="946" height="626" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep1.png" alt="" class="wp-image-1237" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep1.png 946w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep1-300x199.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep1-768x508.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep1-91x60.png 91w" sizes="auto, (max-width: 946px) 100vw, 946px" /><figcaption class="wp-element-caption">Private Endpoint Creation. Step 1</figcaption></figure>



<p>The second step requires a bit more of information, here we will define which resource we try to target with our private endpoint. As expected we need to find our AzureSQL Server here. We fill the combo boxes as usual</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="920" height="529" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep2.png" alt="" class="wp-image-1238" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep2.png 920w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep2-300x173.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep2-768x442.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep2-330x190.png 330w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep2-104x60.png 104w" sizes="auto, (max-width: 920px) 100vw, 920px" /><figcaption class="wp-element-caption">Private Endpoint Creation. Step 2</figcaption></figure>



<p>The third screen is the most important one. We need to select the VNET and Subnets that will host our private endpoint. In this case we want to use databricks so we need to use the VNET we created for databrickks, and then the subnet we created specifically to host the private endpoints.</p>



<p>Another important step here is to integrate it with the DNS. If we dont integrate it, when we use the AzureSQL hostname provided by azure we will still access through the public endpoint. By integrating it in the DNS. the dns queries over the public endpoint in this private zone, will resolve to the private IP of the NIC of the Private Endpoint</p>



<p>If we chose no for the DNS integration then we will have to add static entries in the /etc/hosts or somewhat, or use the private IP instead of the hostname when connecting to the AzureSQL server. To simplify we choose to integrate it.</p>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1024" height="583" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep3-1024x583.png" alt="" class="wp-image-1239" style="width:840px;height:478px" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep3-1024x583.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep3-300x171.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep3-768x438.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep3-105x60.png 105w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep3.png 1241w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption"> Private Endpoint Creation. Step 3.</figcaption></figure>



<p>Once created we should see the private endpoint available. If you look at the right its implemented though a NIC (Network Interface card), and by clicking on it, we can find it and see the ip address assigned:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="177" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep5-1024x177.png" alt="" class="wp-image-1241" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep5-1024x177.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep5-300x52.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep5-768x132.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep5-1536x265.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep5-2048x353.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep5-348x60.png 348w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Our newly created Private Endpoint for Azure SQL</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="225" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep6-1024x225.png" alt="" class="wp-image-1242" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep6-1024x225.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep6-300x66.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep6-768x168.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep6-1536x337.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep6-2048x449.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep6-274x60.png 274w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Finding the PrivateIP Address of the NIC that implements the private endpoint</figcaption></figure>



<h2 class="wp-block-heading">Test the AzureSQL DB Endpoint from Databricks</h2>



<p>Now we have it ready. We can still see from outside that VNET, that our old server still resolves to a public ip, as it was the case before even inside databricks. We can ping it for testing purposes:</p>



<pre class="wp-block-code"><code>C:\Users\Albert&gt;ping azure-sql-server-albert.database.windows.net

Haciendo ping a cr4.westeurope1-a.control.database.windows.net &#91;<strong>104.40.168.105</strong>] con 32 bytes de datos:</code></pre>



<p>As you can see we have a public ip, but lets try to ping it inside the cluster:</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="885" height="178" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep7.png" alt="" class="wp-image-1244" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep7.png 885w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep7-300x60.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep7-768x154.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep7-298x60.png 298w" sizes="auto, (max-width: 885px) 100vw, 885px" /><figcaption class="wp-element-caption">Private endpoint with the DNS integration working fine. Our dns record for the AzureSQL Db does not resolve to a public ip anymore but to the private IP of the PrivateEndpoint</figcaption></figure>



<p>So it&#8217;s working. Its using the private ip instead of the public one. Our last step is to see if we can fetch the data from the database:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="363" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep8-1024x363.png" alt="" class="wp-image-1245" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep8-1024x363.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep8-300x106.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep8-768x272.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep8-1536x544.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep8-2048x726.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep8-169x60.png 169w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Accessing AzureSQL Database though a private endpoint from databricks</figcaption></figure>



<h2 class="wp-block-heading">Creating an AzureDataLake PrivateEndpoint and saving our data to the DataLake through it.</h2>



<p>We are not done yet! We can still complicate matters and create a private endpoint as well to save data to our datalake.</p>



<p>I&#8217;ve created an ADLS Gen 2 storage account, and going back to databricks I see by default it&#8217;s using public access:</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="836" height="194" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep9.png" alt="" class="wp-image-1246" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep9.png 836w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep9-300x70.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep9-768x178.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep9-259x60.png 259w" sizes="auto, (max-width: 836px) 100vw, 836px" /><figcaption class="wp-element-caption">Datalake public access</figcaption></figure>



<p>But we can implement a Private Endpoint as well, and route all the traffic through the azure datacenter itself. Lets see how to do it. For achieving this, we go to our ADS Gen2 storage account, and on the left we click again in Networking, and the second tab is called Private Endpoint connections. We click the plus button to create a new one, and basically we follow the same steps as before with a subtle difference</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="930" height="760" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep10.png" alt="" class="wp-image-1247" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep10.png 930w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep10-300x245.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep10-768x628.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep10-73x60.png 73w" sizes="auto, (max-width: 930px) 100vw, 930px" /><figcaption class="wp-element-caption">Creation of a private endpoint for an ADLS Gen2 storage account.</figcaption></figure>



<p>The difference with a storage account is that we need to chose which api we want to create the private endpoint for. We can use the blob, the table, the queue, the file share and the dfs (DataLake) endpoint (And also the static website!).</p>



<p>We will use the dfs endpoint, and again we will place it in the Private Endpoint subnet of our Databricks vnet. Something like this:</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="814" height="990" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep11.png" alt="" class="wp-image-1248" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep11.png 814w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep11-247x300.png 247w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep11-768x934.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep11-49x60.png 49w" sizes="auto, (max-width: 814px) 100vw, 814px" /><figcaption class="wp-element-caption">Creating a Private Endpoint for our DataLake an dplacing it in the appropiate subnet</figcaption></figure>



<p>After a few minutes our private endpoint will be ready to be used. We can go again to see the NIC and check the private ip or go directly to databricks and ping the storage account url to see if now it&#8217;s resolving to our private endpoint:</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="796" height="184" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep12.png" alt="" class="wp-image-1249" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep12.png 796w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep12-300x69.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep12-768x178.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep12-260x60.png 260w" sizes="auto, (max-width: 796px) 100vw, 796px" /><figcaption class="wp-element-caption">As we can see now databricks resolves our storage account through the private endpoint</figcaption></figure>



<h2 class="wp-block-heading">Test the ADLS Gen2 SA endpoint from Databricks </h2>



<p>If we have the IAM credential Passthrough enabled in our cluster and we have permisison to write to the datalake, now we should be able to write there without going through the internet:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="175" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep13-1024x175.png" alt="" class="wp-image-1250" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep13-1024x175.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep13-300x51.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep13-768x132.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep13-1536x263.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep13-2048x351.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQlPep13-350x60.png 350w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Writing to a DataLake through the Private Endpoint we just created</figcaption></figure>



<p>So this is the end of the tutorial. We created two private endpoints, one for AzureSQL Database and Another for our DataLake and used bot them from Databricks. We also confirmed we are effectively using them by pinging the hostnames of both resources and seeing a change from the public ip to the private one.</p>



<p>Happy Data pojects!</p>
<p>The post <a href="https://www.albertnogues.com/using-azure-private-endpoints-with-databricks/">Using Azure Private Endpoints with Databricks</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Databricks connectivity to Azure SQL / SQL Server</title>
		<link>https://www.albertnogues.com/databricks-connectivity-to-azure-sql-sql-server/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=databricks-connectivity-to-azure-sql-sql-server</link>
		
		<dc:creator><![CDATA[Albert]]></dc:creator>
		<pubDate>Thu, 09 Dec 2021 10:45:34 +0000</pubDate>
				<category><![CDATA[Azure]]></category>
		<category><![CDATA[Databricks]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Spark]]></category>
		<category><![CDATA[SQL]]></category>
		<guid isPermaLink="false">https://www.albertnogues.com/?p=1224</guid>

					<description><![CDATA[<p>Most of the developments I see inside databricks rely on fetching or writing data to some sort of Database. Usually the preferred method for this is though the use of jdbc driver, as most databases offer some sort of jdbc driver. In some cases, though, its also possible to use some spark optimized driver. This &#8230; </p>
<p>The post <a href="https://www.albertnogues.com/databricks-connectivity-to-azure-sql-sql-server/">Databricks connectivity to Azure SQL / SQL Server</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p>Most of the developments I see inside databricks rely on fetching or writing data to some sort of Database.</p>



<p>Usually the preferred method for this is though the use of jdbc driver, as most databases offer some sort of jdbc driver.</p>



<p>In some cases, though, its also possible to use some spark optimized driver. This is the case in Azure SQL / SQL Server. We have still the option to use the standard jdbc driver (what most people do because it&#8217;s standard to all databases) but we can improve the performance by using a specific spark driver. Till some time ago it was only supported with the Scala API but now it&#8217;s possible to be used in Python and R as well, so there is no reason not to give it a try.</p>



<p>In this article we will see the two options to make this connectivity. For the test purposes we will connect to an Azure SQL in the same region (West Europe).</p>



<h2 class="wp-block-heading">Connecting to AzureSQL through  jdbc driver. </h2>



<p>In this case the jdbc driver is already shipped in the databricks cluster, we do not need to install anything. We just can connect directly. Lets see how (We have a scala example <a href="https://docs.microsoft.com/es-es/azure/databricks/data/data-sources/sql-databases" target="_blank" rel="noreferrer noopener">here</a> but i will use python for this example)</p>



<pre class="wp-block-code"><code>#In a real development this should be fetched from a keyvault using a secret scope with: dbutils.secrets.get(scope = "sql_db", key = "username") and  dbutils.secrets.get(scope = "sql_db", key = "password")

jdbcDF = spark.read.format("jdbc") \
    .option("url", f"jdbc:sqlserver://azure-sql-server-albert.database.windows.net:1433;databaseName=databricksdata") \
    .option("dbtable", "SalesLT.Product") \
    .option("user", "anogues") \
    .option("password", "XXXXXX") \
    .option("driver", "com.microsoft.sqlserver.jdbc.SQLServerDriver") \
    .load()

jdbcDF.show()</code></pre>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="362" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR1-1024x362.png" alt="" class="wp-image-1226" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR1-1024x362.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR1-300x106.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR1-768x272.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR1-1536x543.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR1-2048x724.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR1-170x60.png 170w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption>Spark Dataframe from a JDBC Azure SQL DB Source</figcaption></figure>



<p>So as we saw we have been able to connect successfully to our Azure SQL DB using the jdbc driver shipped with databricks. Lets now try to change to the spark optimized driver</p>



<h2 class="wp-block-heading">Connecting to AzureSQL through the spark optimized driver</h2>



<p>To connect using the spark optimized driver, first we need to install the driver in the cluster, as it&#8217;s not available by default.</p>



<p>The driver is available in Maven for both spark 2.X and 3.X. In the microsoft <a href="https://docs.microsoft.com/en-us/sql/connect/spark/connector?view=sql-server-ver15">website</a> we can find more information on where to get them and how to use them. For this exercise purposes we will inbstall it through databricks libraries, using maven. Just add in the coordinates box the following: com.microsoft.azure:spark-mssql-connector_2.12:1.2.0 as can be seen in the image below</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="323" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR2-1024x323.png" alt="" class="wp-image-1227" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR2-1024x323.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR2-300x95.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR2-768x242.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR2-1536x485.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR2-190x60.png 190w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR2.png 2006w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption>Installing the spark AzureSQL Driver from Maven</figcaption></figure>



<p>Once installed we should see a green dot next to the driver, and this will mean the driver is ready to be used. We go back to our notebook and try</p>



<pre class="wp-block-code"><code>#In a real development this should be fetched from a keyvault using a secret scope with: dbutils.secrets.get(scope = "sql_db", key = "username") and  dbutils.secrets.get(scope = "sql_db", key = "password")
jdbcDF = spark.read.format("com.microsoft.sqlserver.jdbc.spark") \
    .option("url", f"jdbc:sqlserver://azure-sql-server-albert.database.windows.net:1433;databaseName=databricksdata") \
    .option("dbtable", "SalesLT.Product") \
    .option("user", "anogues") \
    .option("password", "XXXXXX") \
    .load()

jdbcDF.show()</code></pre>



<p>If we see an error like java.lang.ClassNotFoundException: com.microsoft.sqlserver.jdbc.spark this means that the driver can&#8217;t be found, so probably it&#8217;s not properly installed. Check back the libraries in the cluster and make sure the status is installed. If all goes well we should see again our dataframe:</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="351" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR3-1-1024x351.png" alt="" class="wp-image-1229" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR3-1-1024x351.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR3-1-300x103.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR3-1-768x263.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR3-1-1536x526.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR3-1-2048x701.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR3-1-175x60.png 175w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption> Spark Dataframe from a Spark Azure SQL DB Source </figcaption></figure>



<p>The reason why we should use the optimized spark driver is usually because of performance reasons. Microsoft claims its about 15x faster than the jdbc one. But there is more. The spark driver also allows AAD authentication either by using a service principal or an AAD account, apart of course from the native sql server authentication. Lets try if it works with an AAD account:</p>



<pre class="wp-block-code"><code>jdbcDF = spark.read \
    .format("com.microsoft.sqlserver.jdbc.spark") \
    .option("url", f"jdbc:sqlserver://azure-sql-server-albert.database.windows.net:1433;databaseName=databricksdata") \
    .option("dbtable", "SalesLT.Product") \
    .option("authentication", "ActiveDirectoryPassword") \
    .option("user", "sqluser@anogues4hotmail.onmicrosoft.com") \
    .option("password", "XXXXXX") \
    .option("encrypt", "true") \
    .option("hostNameInCertificate", "*.database.windows.net") \
    .load()
jdbcDF.show()</code></pre>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="366" src="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR4-1024x366.png" alt="" class="wp-image-1230" srcset="https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR4-1024x366.png 1024w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR4-300x107.png 300w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR4-768x274.png 768w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR4-1536x549.png 1536w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR4-2048x732.png 2048w, https://www.albertnogues.com/wp-content/uploads/2021/12/SQLDBR4-168x60.png 168w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<p>For using a service principal you need to generate a token. In python this can be accomplished with the <a href="https://pypi.org/project/adal/" target="_blank" rel="noreferrer noopener">adal</a> library (That needs to be installed in the cluster as well from pypi). You have a sample notebook in microsoft spark driver github account <a href="https://github.com/microsoft/sql-spark-connector/tree/master/samples/Databricks-AzureSQL/DatabricksNotebooks">here</a>.</p>



<p>More information about the driver can be found on the microsoft github repository <a href="https://github.com/microsoft/sql-spark-connector" target="_blank" rel="noreferrer noopener">here</a>.</p>
<p>The post <a href="https://www.albertnogues.com/databricks-connectivity-to-azure-sql-sql-server/">Databricks connectivity to Azure SQL / SQL Server</a> appeared first on <a href="https://www.albertnogues.com">Albert Nogués</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
