Jilles Vreeken
Independent Research Group Leader (W2)
Exploratory Data Analysis
Cluster of Excellence MMCI
Saarland University
Senior Researcher
Databases and Information Systems
Max Planck Institute for Informatics
Saarland Informatics Campus
Building E 1.7 Room 3.22
66123 Saarbrücken, Germany
jilles@mpi-inf.mpg.de
+49 681 302 71 925
Jilles at work in Vancouver

Since October 2013, I lead the independent research group on Exploratory Data Analysis at the DFG cluster-of-excellence on Multimodal Computing and Interaction at the University of Saarland. In addition, I'm affiliated as
Senior Researcher with the Database and Information Systems (D5) group of the Max Planck Institute for Informatics.

My research is mainly concerned with exploratory data mining. That is, I develop theory and algorithms for answering the question `this is my data, tell me what I need to know'. To identify what you need to know, i.e., what is the most interesting structure in the data, I often employ well-founded statistical methods. In particular, Information Theory — the principles of Minimum Description Length (MDL) and Maximum Entropy have proven to be highly valuable tools. Next, I develop highly efficient algorithms for extracting these interesting structures, i.e., models, from very large and complex data—as well as investigate how we can use these structures in a wide range of applications, including identifying rare diseases, e-health, bio-informatics, market analysis, product recommendation, etc.


I'm always looking for talented and motivated PhD candidates, postdocs, and HiWi's
with a strong background in data mining, machine learning, statistics, and/or mathematics.


Currently I'm investigating techniques for identifying informative local structures in large collections of complex data; how to efficiently mine good data descriptions directly such data; the theoretical and practical foundations of interactive exploration of very large data, discovering things by serendipity; how to mine large relational databases; how to mine very large graphs, including characterising influence propagation in social networks; as well as to study well-founded approaches for meaningfully comparing between, and validation of, explorative results.

Below, you'll find an overview of my activities, as well as a selection of my recent publications. You might further be interested in my publications, implementations, our workshop on Interactive Data Exploration and Analytics (IDEA) at KDD'17, or our tutorials on Information Theoretic Methods in Data Mining at ECML PKDD'14 and SIAM SDM'15.


or, in case you're looking for a bit of procrastination, consider
Research in Progress — the secret life of research, through the medium of animated GIFs.


Activities more ▾

Teaching and Advising more ▾
  • Researchers and Assistants
    • Dr. Mario Boley
    • Kailash Budhathoki
    • Sebastian Dalleiger
    • Janis Kalofolias
    • Panagiotis Mandros
    • Alexander Marx
    • Maha Aburahma
    • Iva Baykova
    • Yuliia Brendel
    • Robin Burghartz
    • Simina Ana Cotop
    • Tatiana Dembelova
    • Maike Eissfeller
    • Jonas Fischer
    • Patrick Ferber
    • Magnus Halbe
    • Benjamin Hättasch
    • Michael A. Hedderich
    • Frauke Hinrichs
    • Henrik Jilke
    • David Kaltenpoth
    • Stephanie Lund
  • Former MSc Thesis Students
    • Amirhossein Baradaranshahroudi (2016)
    • Apratim Bhattacharyya (2016)
    • Beata Wójciak (2016)
    • Margarita Salyaeva (2016)
    • Manan Gandhi (2016)
    • Kathrin Grosse (2016)
    • Kailash Budhathoki (2015)
    • Panagiotis Mandros (2015)
    • Thomas Van Brussel (2012)
    • Tanja Van den Eede (2011)
    • Sandy Moens (2010)
    • Andie Similon (2010)
    • Sander Schuckmann (2008)
  • Former BSc Students
    • Frauke Hinrichs (2017)
    • Magnus Halbe (2016)
    • Stefan Bier (2014)
  • Former Research Assistants
    • Shweta Mahajan
    • Sinan Bozca
    • Cristian Caloian
    • Eustace Ebhotemhen
    • Andrea Fuksova
    • Shilpa Garg
    • Tobias Heinen
    • Stefan Neumann
    • Michael Wessely
    • David Ziegler

Selected Recent Publications (go here for the complete list)
In Press
Budhathoki, K & Vreeken, J Origo: Causal Inference by Compression. Knowledge and Information Systems, Springer (IF 2.004)
Fischer, AK, Vreeken, J & Klakow, D Beyond Pairwise Similarity: Quantifying and Characterizing Linguistic Similarity between Groups of Languages by MDL. ComputaciĆ³n y Sistemas (Special Issue for the 18th International Conference on Intelligent Text Processing and Computational Linguistics, CICLing'17)
2017
Budhathoki, K & Vreeken, J MDL for Causal Inference on Discrete Data. In: Proceedings of the IEEE International Conference on Data Mining (ICDM'17), IEEE, 2017 (19.9% acceptance rate).
Marx, A & Vreeken, J Telling Cause from Effect by MDL-based Local and Global Regression. In: Proceedings of the IEEE International Conference on Data Mining (ICDM'17), IEEE, 2017 (full paper, 9.3% acceptance rate; overall 19.9%).
Kalofolias, J, Boley, M & Vreeken, J Efficiently Discovering Locally Exceptional yet Globally Representative Subgroups. In: Proceedings of the IEEE International Conference on Data Mining (ICDM'17), IEEE, 2017 (full paper, 9.3% acceptance rate; overall 19.9%).
Mandros, P, Boley, M & Vreeken, J Discovering Reliable Approximate Functional Dependencies. In: Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), pp 355-363, ACM, 2017 (oral presentation, 8.6% acceptance rate; overall 17.5%).
Bertens, R, Vreeken, J & Siebes, A Efficiently Discovering Unexpected Pattern-Co-Occurrences. In: Proceedings of the SIAM International Conference on Data Mining (SDM), pp 126-134, SIAM, 2017 (25% acceptance rate).
Bhattacharyya, A & Vreeken, J Efficiently Summarising Event Sequences with Rich Interleaving Patterns. In: Proceedings of the SIAM Conference on Data Mining (SDM), pp 795-803, SIAM, 2017 (selected in the top 10 papers of SDM'17, 2.7% acceptance rate; overall 25%).
Budhathoki, K & Vreeken, J Correlation by Compression. In: Proceedings of the SIAM Conference on Data Mining (SDM), pp 525-533, SIAM, 2017 (25% acceptance rate).
Pienta, R, Kahng, M, Lin, Z, Vreeken, J, Talukdar, P, Abello, J, Parameswaran, G & Chau, DH Adaptive Local Exploration of Large Graphs. In: Proceedings of the SIAM International Conference on Data Mining (SDM), pp 597-605, SIAM, 2017 (25% acceptance rate).
Boley, M, Goldsmith, BR, Ghiringhelli, LM & Vreeken, J Identifying Consistent Statements about Numerical Data with Dispersion-Corrected Subgroup Discovery. Data Mining and Knowledge Discovery vol.31(5), pp 1391-1418, Springer, 2017. (IF 3.160) (ECML PKDD'17 Journal Track)
Goldsmith, B, Boley, M, Vreeken, J, Scheffler, M & Ghiringhelli, L Uncovering Structure-Property Relationships of Materials by Subgroup Discovery. New Journal of Physics vol.19, IOP Publishing Ltd and Deutsche Physikalische Gesellschaft, 2017. (IF 3.57)
2016
Budhathoki, K & Vreeken, J Causal Inference by Compression. In: Proceedings of the IEEE International Conference on Data Mining (ICDM'16), IEEE, 2016 (full paper, 8.5% acceptance rate; overall 19.6%). (invited for the KAIS Special Issue on the Best of IEEE ICDM 2016)
Bertens, R, Vreeken, J & Siebes, A Keeping it Short and Simple: Summarising Complex Event Sequences with Multivariate Patterns. In: Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD'16), pp 735-744, ACM, 2016 (oral presentation, 8.9% acceptance rate; overall 18.1%).video
Rozenshtein, P, Gionis, A, Prakash, BA & Vreeken, J Reconstructing an Epidemic over Time. In: Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), pp 1835-1844, ACM, 2016 (18.1% acceptance rate).
Nguyen, H-V & Vreeken, J Flexibly Mining Better Subgroups. In: Proceedings of the SIAM International Conference on Data Mining (SDM), pp 585-593, SIAM, 2016 (overall 25% acceptance rate).
Nguyen, H-V, Mandros, P & Vreeken, J Universal Dependency Analysis. In: Proceedings of the SIAM International Conference on Data Mining (SDM), pp 792-800, SIAM, 2016 (overall 25% acceptance rate).
Nguyen, H-V & Vreeken, J Linear-time Detection of Non-Linear Changes in Massively High Dimensional Time Series. In: Proceedings of the SIAM International Conference on Data Mining (SDM), pp 828-836, SIAM, 2016 (overall 25% acceptance rate).
Athukorala, K, Glowacka, D, Jacucci, G, Oulasvirta, A & Vreeken, J Is Exploratory Search Different? A Comparison of Information Search Behavior for Exploratory and Lookup Tasks. Journal of the Association for Information Science and Technology (JASIST) vol.67(11), pp 2635-2651, Wiley, 2016. (IF 2.26)