Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems

A.M.A. Balayn; C. Lofi; G.J.P.M. Houben

doi:10.1007/s00778-021-00671-8

Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems

A.M.A. Balayn, C. Lofi, G.J.P.M. Houben

Web Information Systems

Research output: Contribution to journal › Article › Scientific › peer-review

24 Citations (Scopus)

334 Downloads (Pure)

Abstract

The increasing use of data-driven decision support systems in industry and governments is accompanied by the discovery of a plethora of bias and unfairness issues in the outputs of these systems. Multiple computer science communities, and especially machine learning, have started to tackle this problem, often developing algorithmic solutions to mitigate biases to obtain fairer outputs. However, one of the core underlying causes for unfairness is bias in training data which is not fully covered by such approaches. Especially, bias in data is not yet a central topic in data engineering and management research. We survey research on bias and unfairness in several computer science domains, distinguishing between data management publications and other domains. This covers the creation of fairness metrics, fairness identification, and mitigation methods, software engineering approaches and biases in crowdsourcing activities. We identify relevant research gaps and show which data management activities could be repurposed to handle biases and which ones might reinforce such biases. In the second part, we argue for a novel data-centered approach overcoming the limitations of current algorithmic-centered methods. This approach focuses on eliciting and enforcing fairness requirements and constraints on data that systems are trained, validated, and used on. We argue for the need to extend database management systems to handle such constraints and mitigation methods. We discuss the associated future research directions regarding algorithms, formalization, modelling, users, and systems.

Original language	English
Pages (from-to)	739-768
Number of pages	30
Journal	The VLDB Journal
Volume	30
Issue number	5
DOIs	https://doi.org/10.1007/s00778-021-00671-8
Publication status	Published - 2021

Keywords

Bias and unfairness
Bias constraints for DBMS
Bias mitigation
Data curation
Decision support systems

Access to Document

10.1007/s00778-021-00671-8Licence: Unspecified

Balayn2021_Article_ManagingBiasAndUnfairnessInDatFinal published version, 1.42 MBLicence: CC BY

Fingerprint

Dive into the research topics of 'Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems'. Together they form a unique fingerprint.

Cite this

@article{2870a84132cf4fbb8cf26ff60529fffc,

title = "Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems",

abstract = "The increasing use of data-driven decision support systems in industry and governments is accompanied by the discovery of a plethora of bias and unfairness issues in the outputs of these systems. Multiple computer science communities, and especially machine learning, have started to tackle this problem, often developing algorithmic solutions to mitigate biases to obtain fairer outputs. However, one of the core underlying causes for unfairness is bias in training data which is not fully covered by such approaches. Especially, bias in data is not yet a central topic in data engineering and management research. We survey research on bias and unfairness in several computer science domains, distinguishing between data management publications and other domains. This covers the creation of fairness metrics, fairness identification, and mitigation methods, software engineering approaches and biases in crowdsourcing activities. We identify relevant research gaps and show which data management activities could be repurposed to handle biases and which ones might reinforce such biases. In the second part, we argue for a novel data-centered approach overcoming the limitations of current algorithmic-centered methods. This approach focuses on eliciting and enforcing fairness requirements and constraints on data that systems are trained, validated, and used on. We argue for the need to extend database management systems to handle such constraints and mitigation methods. We discuss the associated future research directions regarding algorithms, formalization, modelling, users, and systems.",

keywords = "Bias and unfairness, Bias constraints for DBMS, Bias mitigation, Data curation, Decision support systems",

author = "A.M.A. Balayn and C. Lofi and G.J.P.M. Houben",

year = "2021",

doi = "10.1007/s00778-021-00671-8",

language = "English",

volume = "30",

pages = "739--768",

journal = "The VLDB Journal",

issn = "1066-8888",

publisher = "Springer",

number = "5",

}

Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems. / Balayn, A.M.A.; Lofi, C.; Houben, G.J.P.M.
In: The VLDB Journal, Vol. 30, No. 5, 2021, p. 739-768.

Research output: Contribution to journal › Article › Scientific › peer-review

TY - JOUR

T1 - Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems

AU - Balayn, A.M.A.

AU - Lofi, C.

AU - Houben, G.J.P.M.

PY - 2021

Y1 - 2021

N2 - The increasing use of data-driven decision support systems in industry and governments is accompanied by the discovery of a plethora of bias and unfairness issues in the outputs of these systems. Multiple computer science communities, and especially machine learning, have started to tackle this problem, often developing algorithmic solutions to mitigate biases to obtain fairer outputs. However, one of the core underlying causes for unfairness is bias in training data which is not fully covered by such approaches. Especially, bias in data is not yet a central topic in data engineering and management research. We survey research on bias and unfairness in several computer science domains, distinguishing between data management publications and other domains. This covers the creation of fairness metrics, fairness identification, and mitigation methods, software engineering approaches and biases in crowdsourcing activities. We identify relevant research gaps and show which data management activities could be repurposed to handle biases and which ones might reinforce such biases. In the second part, we argue for a novel data-centered approach overcoming the limitations of current algorithmic-centered methods. This approach focuses on eliciting and enforcing fairness requirements and constraints on data that systems are trained, validated, and used on. We argue for the need to extend database management systems to handle such constraints and mitigation methods. We discuss the associated future research directions regarding algorithms, formalization, modelling, users, and systems.

AB - The increasing use of data-driven decision support systems in industry and governments is accompanied by the discovery of a plethora of bias and unfairness issues in the outputs of these systems. Multiple computer science communities, and especially machine learning, have started to tackle this problem, often developing algorithmic solutions to mitigate biases to obtain fairer outputs. However, one of the core underlying causes for unfairness is bias in training data which is not fully covered by such approaches. Especially, bias in data is not yet a central topic in data engineering and management research. We survey research on bias and unfairness in several computer science domains, distinguishing between data management publications and other domains. This covers the creation of fairness metrics, fairness identification, and mitigation methods, software engineering approaches and biases in crowdsourcing activities. We identify relevant research gaps and show which data management activities could be repurposed to handle biases and which ones might reinforce such biases. In the second part, we argue for a novel data-centered approach overcoming the limitations of current algorithmic-centered methods. This approach focuses on eliciting and enforcing fairness requirements and constraints on data that systems are trained, validated, and used on. We argue for the need to extend database management systems to handle such constraints and mitigation methods. We discuss the associated future research directions regarding algorithms, formalization, modelling, users, and systems.

KW - Bias and unfairness

KW - Bias constraints for DBMS

KW - Bias mitigation

KW - Data curation

KW - Decision support systems

UR - http://www.scopus.com/inward/record.url?scp=85105874163&partnerID=8YFLogxK

U2 - 10.1007/s00778-021-00671-8

DO - 10.1007/s00778-021-00671-8

M3 - Article

SN - 1066-8888

VL - 30

SP - 739

EP - 768

JO - The VLDB Journal

JF - The VLDB Journal

IS - 5

ER -

Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this