官术网_书友最值得收藏!

Engineering data versus model variety

Having a large choice of algorithms for your predictions is always a good thing, but at the end of the day, domain knowledge and the ability to extract meaningful features from clean data is often what wins the game.

Kaggle is a well-known platform for predictive analytics competitions, where the best data scientists across the world compete to make predictions on complex datasets. In these predictive competitions, gaining a few decimals on your prediction score is what makes the difference between earning the prize or being just an extra line on the public leaderboard among thousands of other competitors. One thing Kagglers quickly learn is that choosing and tuning the model is only half the battle. Feature extraction or how to extract relevant predictors from the dataset is often the key to winning the competition.

In real life, when working on business related problems, the quality of the data processing phase and the ability to extract meaningful signal out of raw data is the most important and time consuming part of building an efficient predictive model. It is well know that "data preparation accounts for about 80% of the work of data scientists" (http://www.forbes.com/sites/gilpress/2016/03/23/data-preparation-most-time-consuming-least-enjoyable-data-science-task-survey-says/). Model selection and algorithm optimization remains an important part of the work but is often not the deciding factor when implementation is concerned.

A solid and robust implementation that is easy to maintain and connects to your ecosystem seamlessly is often preferred to an overly complex model developed and coded in-house, especially when the scripted model only produces small gains when compared to a service based implementation.

主站蜘蛛池模板: 天津市| 伽师县| 阿巴嘎旗| 新宁县| 榕江县| 五家渠市| 盐津县| 宁陵县| 白山市| 道孚县| 咸丰县| 彩票| 六枝特区| 罗定市| 临朐县| 定州市| 固镇县| 延安市| 青阳县| 佛坪县| 敖汉旗| 科技| 巢湖市| 南澳县| 保靖县| 赤水市| 遂昌县| 明溪县| 武乡县| 柘城县| 大埔县| 麻栗坡县| 西藏| 集安市| 普洱| 虎林市| 玉屏| 屯留县| 乌恰县| 廊坊市| 绥阳县|