Skip to main content

Nimble Semantic Web Search on Databricks

Summary

  • Nimble integrates semantic web search directly into the Databricks Data and AI platform, organizing live web data into task-specific indexes with entities, fields, relationships, and freshness context before reasoning begins — rather than returning raw pages from a generic index.
  • AI agents using Nimble can combine internal lakehouse data with real-time external web signals inside Databricks, governed by Unity AI Gateway, with Nimble benchmarks showing 30–50% accuracy gains over generic web search for enterprise use cases.
  • Nimble is deployed with Databricks Lakebase, integrates with Genie, Genie Code, and MCP, and is available through the Databricks Marketplace, enabling enterprises to add live web intelligence to their agents and analytics workflows without leaving the governed lakehouse environment.

Nimble Semantic Web Search on Databricks

Watch: Nimble Semantic Web Search on Databricks
Nimble brings semantic web search into Databricks by turning high-volume web search and real time crawling into task-specific web indexes. Instead of giving agents raw pages from a generic index, Nimble organizes web data into entities, fields, relationships, freshness, and source context before reasoning begins. Deployed with Databricks Lakebase and aligned with Unity AI Gateway, Nimble lets enterprises consume live web intelligence directly inside their governed lakehouse environment. This enables AI applications that combine internal data with fresh external web context, securely and at scale. We will show how a Fortune 500 company deploys a task-specific agent with Nimble’s specialized web search, delivering more accurate web intelligence with lower token cost and enterprise-scale reliability.
Talk By: Uri Korovich, CEO & Co founder, Nimble ;

Chapters

FAQs

What is semantic web search and how is it different from generic web search?

Semantic web search, as described in this video, organizes web data into structured entities, fields, relationships, freshness scores, and source context before an agent reasons over it, rather than returning a ranked list of raw web pages. This means agents receive contextualized, structured intelligence that is immediately useful for analysis rather than having to parse unstructured HTML. Nimble benchmarks show this task-specific approach delivers 30–50% accuracy gains over generic web search results for enterprise use cases.

Why do AI agents need live web data alongside internal lakehouse data?

Internal lakehouse data captures transactions, CRM records, and operational history, but it has no visibility into what is happening outside the organization — social media signals, competitor activity, regulatory changes, and market events that directly affect business outcomes. This video illustrates the gap with a consumer goods example where a sales decline traced back to a social media signal that existed only in external data. Combining internal and external data through Nimble on the Databricks Data and AI platform gives agents a complete picture for analysis and autonomous decision-making.

How does Nimble integrate with Databricks?

Nimble is deployed with Databricks Lakebase as the data layer and aligns with Unity AI Gateway for governance and access control over all web data queries. It integrates natively with Genie, Genie Code, and MCP, so agents and end users can query Nimble's web-enriched tables directly within the Databricks Data and AI platform without switching to a separate tool. This video includes a live demo showing Nimble being set up inside Databricks and Genie querying web-enriched tables in real time.

What enterprise use cases benefit most from semantic web search in Databricks?

According to this video, semantic web search is most valuable when internal data alone cannot answer the question — tracking competitor pricing, monitoring social signals affecting demand, identifying regulatory changes, or sourcing external intelligence for go-to-market analysis. Task-specific web indexes are built for domain routing so each agent queries only the relevant portion of the web for its use case, reducing token cost and improving answer precision. A Fortune 500 company is shown deploying a task-specific agent with Nimble's specialized search to deliver more accurate intelligence at enterprise scale and lower cost.

Full transcript

[00:09] Uh, welcome. Nice to meet you everyone. I'm Uri from Nimbell. Um, we are Let's wait one more second. I see everybody just putting their help. Great. So, great to meet you all. I'm Uri. I'm the
[00:21] founder and CEO of Nimbo. I've been um doing data science since forever and started this company about five and a half years ago working with data bricks a lot and I've been working almost all
[00:34] my life on search system and data science in many cases you need to work with your internal data and then you pulling the right analysis for the businesses then you can run everything
[00:46] that you want what we are going to talk about today how nimble is helping everyone's because it's working within data bricks as the data lakehouse to integrate the web to running web search
[00:58] into the data lake. Um for those who don't know Nimble, we are portfolio portfolio company of uh data bricks that invested us uh far back in round B and uh we're working with the links brand.
[01:11] There's not going to be a sales pitch. So the only things I'm going to talk about how web search can transform data science work within data bricks and especially whenever we want to start deploying agents. So let's dive in. Why
[01:24] web search is important. We have all the internal context within our data lakehouse. But the only things that we don't have is that what's happening behind the four
[01:36] walls of our organization. For a really long time, databicks had the concept of marketplace and data marketplace that's helping us to source data from vendors. Um these are great if we know exactly what we want. But in many cases, our
[01:51] questions are open. A lot of things happening outside organization. we sometimes don't even know what exactly we want. So we have a great data inside the data warehouse but the problems that we are facing in many cases that's the
[02:04] data that we are looking is not only the transactions or our CRM the data that we are looking is external signal those working together can give us the complete analysis what we're going to do today I'm going to share few examples of
[02:19] how internal data and external data will working first something that we saw recently in the past few weeks so how I'm getting everything that's happening in my internal organization but it's need to work together with something
[02:32] that's got happening outside. Then we would break for a few minutes and doing a live demo of Niml on datab bricks. How we running the web search and setting it up. And then we're going to finish up with some technical um under the hood
[02:45] how things working, how has it been integrated into the unity catalog and the ontology and all the database ecosystem. So let's start the first examples that we're seeing
[02:57] here and this is live examples working with a fortune 100 um CPG that's recently in the past few weeks they saw a huge spike in their sales the spike in
[03:09] the sales was in specific region and only in the US and basically they started to analyze and using the power of the data warehouse what is the trends that's happening and we're seeing the revenue and the units sold had a great
[03:23] correlation but the question that the business came to the data science team and data analytics team what happened what is the reason and as we know with the business wants that these spikes
[03:36] will continue want to understand why so a lot of urgency became into the organization and one day somebody came in and said okay search for the external informations that can affect all my
[03:50] Arizona sales in the specific region and the beautiful part of that is right now we can ask genies this type of question genie know the internal context of what's happening in my cell so I can say
[04:03] my sales in Arizona you know how to pull the data from the relevant table that we need so we don't need to do all the overhead write a very long notebook what we need to do in the past and then going to the web and starting searching for
[04:15] signals that might be relevant here in that examples it's looking for a search and social and many others And what we have found that in that specific date that's been running is
[04:29] getting a FIFA or sorry a FIFA soccer um player, one of the Argentina um top players been talking about that specific
[04:42] brand, how we love to use their product and then in the same day that's the game was not far away they got the huge spike in sales. So the outcome of that for the people working in a data science data
[04:54] analytics teams that they found the right correlation what happened and now what's happening we just got an update two days ago they've been starting a contract with the agency of the same soccer player to on boarding as an
[05:07] influencers to work with a brand so a classic data science teams is connected to the business something happening in the business and right now I can simply go to Genie ask question search the web
[05:19] and find the relevant information. This is one examples of how the power of data bricks, the power of Genie and the power of all the data warehouse combined together with web search driving
[05:31] analytics for the business. The wall story eventually is boiled down to that simple diagram. In one hand, we have everything that's happening inside the organization and in
[05:45] the second hand we have everything that's happening outside the organization. Having those together create the unified intelligence that's needed to have the combined analysis
[05:57] to unlock the question themselves as a data scientist many cases you come and ask yourself I see that correlation I'm trying to understand how I can forecast everything that happening to my business I have this data set but this data set
[06:11] is missing I need to enrich it with external data or in many cases I've been asking myself I wish I can organize the web in a very nice governed data tables.
[06:24] Yes, the web is messy. The web got a lot of different um data in many different websites. It's very hard to govern. And here we can combine that into the data intelligence platform.
[06:37] So what is web search with semantic layer? something that we have spent a lot of time in the lab to understand how we can not only serve the most relevant links that's happening from the web how
[06:49] not only to create the crawlers that extract the information but also to map the specific entities and the resolutions so we've got the entities we've got the fields we need to understand the relationship the taxonomy
[07:01] and all the source signals that's happening together imagine that I want to get create a table all of my competitor pricing That's a very complex skew mapping and
[07:14] catalog mapping. Everything need to be governed high scale and secure because eventually if you're going to fit in a messy data, no way that we can provide analytics that someone in the business
[07:26] can actually trust. Here's an example of company due diligence pipeline. Eventually what we are going to do the same ways that we're working all the time the same way that databicks is allowing us to create
[07:39] notebooks and workflows that provide the value for the business either that by the end of it I'm going to have a dashboard or an agent I want to build a pipeline that is part of my pipeline I must have an option to get the web
[07:53] search because if not I'm just have everything that's happening inside my organization and the power of how we've been designed the nimble web search system to work in tandem with the unity
[08:05] catalog and write the data also to data bricks. So unlike any other search engines that provide just snippets and results back to the user here in that case we can actually generate tables in
[08:18] the lakehouse. These tables can be in a delta format or JSONs if you're working on lakebase something that's in the heart of web search for data science web search for
[08:30] high completeness high accuracy enterprise use case is data lineage. the exact same way that every pipelines that we are running must have all the parameters that's allowing us to know that we can trust this data. So it's
[08:43] starting with the task definition is going all the way to the domain routing because we're going to find this information in many different domains. There are some domains that we can allow because we're trusting these domains. Some other domains we are not trusting.
[08:56] So we have full visibility everything that's happening in the nimble web search algorithm. After that we're doing a source ranking. We're doing a workflow steps and getting the guardrails. Only by the end of it, we're gonna have the
[09:08] structural data tables. Everything that's happening here on the left side, it's the nimble web search algorithm. So you as a team don't need to care about it at all. Everything is configured by an API. But the outputs that we are
[09:20] getting is a clean structure data table says I can run my business or ground my agent bricks genies. So, I hate whenever we're doing these presentations and we're doing a lot of
[09:33] like slides and talks, but we're all technical people. So, let's see that in action. Um, the first things that I'm doing here is getting into the database marketplace. Um, in the marketplace, the
[09:46] only things that I need to do, I will go here. I will search for Nimbel and I'm going to find Nimble MCP. After I got an email MCP, I'm going to click install here and I need to put a name
[10:00] for that. So that's going to be an email web search and I need to have an API key. So I can simply go to our website. I can sign up directly and after I'm signing up, I will be
[10:13] inside the app and I can create an API key. I will create an API key here. I will call it data bricks. Now I have this key. I will copy this
[10:25] key and I will go to databicks and paste it here. Now I'm up and running and I can run everything in my database environment. What I'm going to show you right now is
[10:38] that's how I'm connecting database Genie code to run some questions. And here in that case I'm asking J genie to compare running shoes and prices deals across different retailers in
[10:51] Amazon and Walmart. Genie is running right now providing the analytics that I want. Genie will process all the data and write SQL on top of the searches that we have got from Nimble. And Genie
[11:04] will finish all the way. And here I'm gonna stop for a second because the next things that I'm going to do, I got all the data from Genie. Now I'm happy, but I want to build something that's going to be meaningful for my business. So
[11:18] Genie Code is such a great product. I love using Geniode because I can move way faster. And basically I'm telling Genieode, I want to create a full-blown analytics on top of this data. I want to create a BII dashboard that I can share
[11:31] with my business. So basically Genie right now plot an AIBI dashboard of best running shoes Amazon versus Walmart powered by Nimble and here I can see all this data average price by source
[11:43] average by keyword and source price by rating so forth and so on. This is so beautiful. I'm showing you a dashboard but we are eventually a data people. So I would do a one step ahead and I will
[11:55] say Genie you know it's really nice that uh I want to show the data and I want to add you to add this data to my unity catalog. So right now Genie find 520
[12:07] rows that we have got from Amazon and we can scale it to thousands and hundred thousands and even millions of rows and Genie created with Nimble web search delta tables within my data warehouse
[12:19] and all these data mapped directly to my Unity catalog. The reason that's important because right now I understand that this web search systems not only create a governed web search that I can
[12:33] rely in my day-to-day business. This web search system is not only works for me whenever I need to ground my agents that are running on database to have the context but I can actually create a
[12:45] fullblown analytics on top of everything that's happened with web search and the power of databicks. So returning back to where we started in the past we have the data marketplace and the only way to get
[12:57] data that's coming on from the web I need to have crawling or data vendors or dealing with very hard work to comply and get good data with web scraping now web search systems and data bricks
[13:09] basically automating all that flow and I can get clean governed trusted source of data directly for my database environment and that's can work across different verticals Here he show
[13:22] examples example of CPG use case to see the pricing but I can monitor everything that's happening with drug discovery on healthcare and life size everything that's happening on compliance of my product. I can track everything that's
[13:36] happening in the public market in financial services insurance companies using Nimble and web search system together to automate claim processing and underwriting because the web is the biggest database.
[13:48] Now when combining the web with datab bricks intelligence platform we can get this value. So how can we integrate web search system directly to databicks? We have
[14:01] three main options. The first one is Genie. We talked about it. I can use Genie code. I can use Genie1. I can use any type of GI integration directly for my data. The second options that we have
[14:13] is databicks agents. database just just launched their agents that can run on top of database. I can create business facing agents. I can create an agent that will do data enrichment for me or
[14:25] we can go all the way with MCP client. In many cases, I prefer not to use GI code directly. I prefer to use my cloud code and connect it to database CLI. I'm getting better results. So, Nimble do
[14:38] have direct integration into cloud. We have actually more than one and a half million users and developers that's download and connected our plug-in to uh cloud to run that um or I can work
[14:50] directly from codex or from copilot. So direct CLI it's very easy. If I'm working more of a classic of a data data science or BI teams we have a UDF so we can call everything from the UDF
[15:03] eventually everything happening to the data warehouse either to create data tables in wrench column and everything got the lineage and the all the trails. So if I need to go all in here and showing uh how easy is to work with
[15:17] neighbors, I can go to the docs in the integration. I have here database on the left side and I have everything that I need to jump start from the MCP to gen1 if I
[15:29] want to create directly and business facing. So I can add the MCP here and I can simply ask any question that's the business directly. They don't need to interact with me. And the beautiful of connecting Nimble directly to Genie1
[15:42] because Genie already been ground on all of the internal data that I have in my organization. So connecting Nimble directly to Genie1 helping me and every business users to ask any questions that's happening on the web with a
[15:55] combine everything that's happening internally. I can do data enrichment assuming that I have a tables and I want to enrich this table. So I can use the web search for that or I can use Genie code the examples that I show before.
[16:08] something that I'm very excited to share today. Um, in the keynotes this morning, databicks announced the omni agents, the new open source technologies that databicks using the meta harness to allow every agent with every LLM to run
[16:22] in the governed web uh in the governed databicks environment. So as you can see here um every CLM or custom agents do have the runner running on the databic server and build this terminal UI where
[16:35] all the applications happening on the database side. Nimble is a launch partner for the Omni agent. So here I can show you an example. This is the Omni agent and you can right now get and clone the GitHub rep of the Omni agents
[16:48] or get it directly in your database environment. And inside omni agent you can ask any question. You can go here. Nimble agent is part of the pick bar. So if I will stop here for a second you
[17:01] will see that you can choose cloud codeex or even add nimble agent as a web search. And right now my cloud code can work directly with omni agents. And I'm asking myself what is the mission and
[17:14] the vision of data bricks and it will go we'll find the relevant information from the web directly from my agents to ground the results so that my agents right now can work directly with the internal data and the external data as
[17:27] well. So this is another example so how we can deploy Nimble on top of databix environment.
[17:39] So I want to finish up with how web search agents work. Eventually we have four layers. There is underlying technologies that's allowing us to provide high accuracy and high completeness web data that trusted in
[17:51] the tables within databicks. The first one is the query grounding. Every queries that we're sending to Nimble is going to be ground with the ontology of our organization. So if someone is going to ask what my competitors new product
[18:05] launches and moves we will know who are the competitors because this data has been aware in my database tables or returning back to the examples whenever I'm asking what is the sales um the last sales drop I know that information
[18:19] retrieval and ranking the algorithm is running on the database environment so we always got the full audit trail and our data never leaving the data warehouse we've got internal context enrichment And eventually we're getting
[18:32] the semantic web search. Nimble under the hood is built on databicks. So whenever we're running our web search and getting all the information from the web, we store the data directly in
[18:44] Nimble databicks environment and then we have the extract the duplication we're doing score mapping and entity resolution building the knowledge graph and then eventually we're processing the data on Nimble databicks environment.
[18:56] Eventually this is helping us to get the high accuracy and high completed that's needed to provide that result but another outcome whenever you're working on data bricks it's super fast because we're doing zero copy and delta sharing
[19:08] on RSI so it's really easy to work with in that sense the values that we are getting with high accuracy high completeness web search coming from
[19:20] using the NL system is 30 to 50% more accurate result comparing to any other web search system that's been out there. The second part which is also important whenever you start deploy agents is 40%
[19:33] less tokens in the overall usage and we need to perform three times less searches to find the relevant information because we have the relevant context that you have inside the data
[19:46] warehouse. Right now every search that you're running is know exactly where to find that information. So the value whenever I'm building an agent either there's going to be a due diligence risk agent or I build a competitive pricing
[20:00] or competitive intelligence complete map view I need to do less searches in the web because nimble know exactly where to search to sum everything up um nimble add your database environment the
[20:13] opportunity to do web search I just show in that simple presentation how we can go to the marketplace install nimble sign up to nimbleway.com get the API I key and start working with that environment. We don't need to go into it
[20:27] and security because it's on the marketplace. Dataware's already done all the security checks that's needed. The power is that what we are seeing and that's something that we're really excited is that data practitioners and business leaders across different
[20:40] verticals build extremely interesting data applications. We have customers in the life science that's going and doing a very deep drug discovery. We have lawyer firms that tracking patents
[20:52] that's going everything on the public web and recently we saw a lot of cyber companies that's checking checking for vulnerabilities um and also the misconfiguration in users manual and the amount of data
[21:04] that's been available on the web is enormous and right now you have the power of taking the data warehouse the data lakehouse all the data that you already have in your organization and enrich it together with external web
[21:18] thank you and I'm really excited to see what you've We're going to build

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.