Rule Discovery for (Semi-)automatic Repairs of ETL Processes
[ 1 ] Instytut Informatyki, Wydział Informatyki i Telekomunikacji, Politechnika Poznańska | [ P ] pracownik
2020
rozdział w monografii naukowej / referat
angielski
- data source evolution
- ETL process repair
- Case-Based-Reasoning
- rule discovery from cases
EN A data source integration layer, commonly called extract-transform-load (ETL), is one of the core components of information systems. It is applicable to standard data warehouse (DW) architectures as well as to data lake (DL) architectures. The ETL layer runs processes that ingest, transform, integrate, and upload data into a DW or DL. The ETL layer is not static, since the data sources being integrated by this layer change their structures. As a consequence, an already deployed ETL process stops working and needs to be re-designed (repaired). Companies typically have deployed from thousands to hundreds of thousands of ETL processes. For this reason, a technique and software support for repairing semi-automatically a failed ETL processes is of vital practical importance. This problem has been only partially solved by technology or research, but the solutions still require an immense work of an ETL administrator. Our solution is based on a case-based-reasoning combined with repair rules. In this paper, we contribute a method for automatic discovery of repair rules from a stored history of repair cases.
12.08.2020
250 - 264
20
70