Databases & Machine Learning

We develop and analyze database foundations for feature engineering

In the design of analytical procedures and machine-learning solutions, a critical and time-consuming task is that of feature engineering, for which various recipes and tooling approaches have been developed. We develop and analyze database foundations for feature engineering, with the goal of opening the way to research and techniques to assist developers by utilizing the database’s modeling and understanding of data and queries, and by deploying the well studied principles of database management.

People

Publications

  1. Accelerating the Global Aggregation of Local Explanations

    Alon Mor, Yonatan Belinkov, Benny Kimelfeld

    AAAI 2024

    Databases & Machine Learning

  2. Selecting Walk Schemes for Database Embedding

    Yuval Lubarsky, Jan Tönshoff, Martin Grohe, Benny Kimelfeld

    CIKM 2023: 1677-1686

    Databases & Machine Learning

  3. Regularizing Conjunctive Features for Classification

    Pablo Barceló, Alexander Baumgartner, Victor Dalmau, Benny Kimelfeld

    PODS 2019: 2-16

    Databases & Machine Learning

  4. A Relational Framework for Classifier Engineering

    Benny Kimelfeld, Christopher Ré

    SIGMOD Record 47(1): 6-13 (2018)

    Databases & Machine Learning

    Abstract

    In the design of analytical procedures and machine-learning solutions, a critical and time-consuming task is that of feature engineering, for which various recipes and tooling approaches have been developed. We embark on the establishment of database foundations for feature engineering. Specifically, we propose a formal framework for classification in the context of a relational database. The goal of this framework is to open the way to research and techniques to assist developers with the task of feature engineering by utilizing the database’s modeling and understanding of data and queries, and by deploying the well studied principles of database management. We demonstrate the usefulness of the framework by formally defining key algorithmic challenges and presenting preliminary complexity results.