Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There's a key statement here that most will either miss or never reach. The second last paragraph

> Machine Learning Automation

Tuning models is a pain, and if you don't understand the model and all its parameters (like most software engineers who's job is the write software and not build models!) it takes time and can be immensely frustrating. Randal Olson[1] just announced TPOT[2], a Python tool that "automatically creates and optimized machine learning pipelines using genetic programming". This is going to be a huge lever for engineers wanting to experiment/implement with ML algorithms.

[1] http://www.randalolson.com/2016/05/08/tpot-a-python-tool-for... [2] https://github.com/rhiever/tpot



Yes, but does it work? I cannot tell from reading the introductory materials whether it will actually search efficiently over large space (where it is computationally expensive to evaluate the fitness of any individual point).

I know that folks are having some success with Gaussian process optimisation for hyper parameter tuning (eg https://github.com/Yelp/MOE), but I would be sceptical about the application of genetic programming to this area, particularly as broadly defined a search space as TPOT seems to set up (genetic programming either doesn't make assumptions about the objective function being optimised which can be used to search more efficiently, like Gaussian optimisation does - or uses implicit assumptions that may be unsuitable, depending on how you want to think about it); has anyone seen benchmarks against other search methods? The workflow does look great though - integrating it into scikit looks really clever.


I agree and I think it's clear TPOT and similar tools, are the first generation. Genetic algorithms might be slow/costly but the concept is there with an integrated, implementable solution. If it's as useful as I think it could be, there will be a flood of modules, libraries and packages developed with efficiency, UI and domain specific improvements.


Genetic algorithms are also immensely slow, perhaps too slow to be useful on a very large dataset.


I'm a bit confused by this idea, maybe someone can clarify. AFAIK, genetic programming/algorithms themselves require some degree of tuning in order to be useful in a reasonable amount of time and for avoiding local minima each generation, via mutation rates, population size, randomized simulated catastrophe, etc. That being said, how does something like TPOT actually ease the pains of tuning the parameters of a different model, if it requires it's own tuning?


There is something similar for Tensorflow:

https://github.com/calvinschmdt/EasyTensorflow

"This package provides users with methods for the automated building, training, and testing of complex neural networks using Google's Tensorflow module. "


Thanks for the link to TPOT, this looks really interesting




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: