Company

Databricks

別名: データブリックス, Databricks

databricks.com

Overview

最終更新: 2026年7月11日

Apache Sparkの創設者らによって設立されたデータ分析プラットフォーム企業。データウェアハウスとデータレイクの利点を組み合わせた「レイクハウス」アーキテクチャを提唱し、企業のデータ活用とAI導入を支援する。

Mentioned Articles

9 件

Research Papers

5 件
  • Cloud and Data Transformation in Banking: Managing Middle and Back Office Operations Using Snowflake and Databricks

    Venugopal Tamraparani

    202514 件引用Semantic Scholar
  • Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads

    Martin Grund, Stefania Leone, Herman van Hövell, Sven Wagner-Boysen, Sebastian Hillig, Hyukjin Kwon, David Lewis, Jakob Mund, Polo-Francois Poli, Lionel Montrieux, Othon Crelier, Xiao Li, Reynold Xin, Matei Zaharia, M. Petropoulos, T. Papathanasiou

    20255 件引用Semantic Scholar

    Today, upgrading Apache Spark versions typically involves significant effort, with unclear investment requirements, including trial and error. This is mainly because in Spark there is no clear separation between the application code and the engine code. Apart from changes in the public API, any Spark internal changes may affect user workloads as users may rely on Spark internals: bug fixes in the Spark engine, changes to internal APIs, library upgrades, or language upgrades may affect customer workloads. As a result, Spark users are often reluctant to upgrade. The downside is that performance improvements, bug fixes, and new features take significantly longer to adopt, preventing customers from quickly benefiting from these improvements. In addition, it increases engineering complexity to manage a large number of different Spark versions. For Databricks serverless jobs and notebooks, we fundamentally transformed and simplified the user experience when using Spark. We shifted user focus from managing Spark runtime versions to managing the stable API that they integrate with - we created client-versioned workloads with a versionless Spark server. Decoupling the client from the Spark engine using Spark Connect has enabled Databricks to automatically upgrade the Spark server, providing users faster access to the latest features while minimizing disruptions from both intentional and unintentional breaking changes, all without compromising workload compatibility and with zero code changes needed from the user. This approach also offers significant benefits to Databricks by streamlining the release process, consolidating usage onto fewer server versions, and reducing engineering overhead from needing to backport changes. In this demonstration, we will first briefly introduce the architectural foundation of versionless Spark, leveraging Databricks' multi-user Spark compute and Spark Connect, followed by describing in more detail how we manage seamless upgrades for our customers, and finally talk about what the user experience is.

  • Gen AI For ELT (Extract, Load, Transfer) in Streaming Application with Databricks/Snow Flakes

    Muhamed Ramees Cheriya Mukkolakkal

    20253 件引用Semantic Scholar

    This paper provides a system literature review of the implementation of Generative Artificial Intelligence (GenAI) in ELT (Extract, Load, Transform) pipelines to incoming applications, concentrating on the Databricks and Snowflake services. The review is based on the summary of the results of fifty chosen studies devoted to the study of GenAI-based automation, scalability and adaptive transformation in real-time data processing. It is shown that GenAI drastically increases the intelligence of the pipeline and its work efficiency and allows working with dynamic schemas and with customised analytics. Nevertheless, data quality, data governance, explainability, and human control are still largely on the agenda. The research suggests a pathway to hybrid ELT architectures to combine GenAI automation and sound governance procedures to establish reliability and responsible execution in the streaming setting.

  • A Secure-by-Design Approach to Big Data Analytics Using Databricks and Format-Preserving Encryption

    Juan Lagos-Obando, Gabriel Aillapán, Julio Fenner-López, Ana Bustamante-Mora, M. Burgos-López

    20252 件引用Semantic Scholar

    Managing and analyzing data in data lakes for big data environments requires robust protocols to ensure security, scalability, and compliance with privacy regulations. The increasing need to process sensitive data emphasizes the relevance of secure-by-design approaches that integrate encryption techniques and governance frameworks to protect personal and confidential information. This study proposes a protocol that combines the capabilities of Databricks and format-preserving encryption to improve data security and accessibility in data lakes without compromising usability or structure. The protocol was developed using a design science methodology, incorporating findings from a systematic literature review and validated through expert feedback and proof-of-concept experiments in banking environments. The proposed solution integrates multiple layers, data ingestion, persistence, access, and consumption, leveraging the processing capabilities of Databricks and format-preserving encryption to enable secure data management and governance. Validation results indicate the protocol is effectiveness in protecting sensitive data, with promising applicability in regulated industries. This work contributes to addressing key challenges in big data security and lays the groundwork for future developments in data governance and encryption techniques.

  • AI-augmented real-time retail analytics with spark and Databricks

    Lingareddy Alva

    20252 件引用Semantic Scholar

    AI-augmented real-time retail analytics represents a transformative approach for modern retail operations, enabling businesses to process and act on data instantaneously in an increasingly competitive landscape. This comprehensive technical article explores the architecture, implementation, and business applications of an integrated analytics platform built on Apache Spark, Databricks, and Azure Event Hubs. The platform ingests data from diverse sources including IoT devices, point-of-sale systems, e-commerce platforms, mobile applications, and social media to create a unified view of retail operations. Advanced machine learning capabilities enable demand forecasting, customer segmentation, price optimization, and fraud detection with unprecedented accuracy. Large language models further enhance the platform by enabling natural language queries and automated insight generation, democratizing access to analytics across retail organizations. The business impact encompasses hyper-personalized customer experiences, predictive inventory management, revenue optimization strategies, and operational efficiency improvements. Implementation considerations and future trends are discussed, providing a blueprint for retailers seeking to leverage real-time analytics as a competitive differentiator in the age of artificial intelligence.

External Mentions

10 件