Live data from Hacker News

The anatomy of an ML-powered stock picking engine

principiamundi.com

71–80 of 109 posts

Re: The anatomy of an ML-powered stock picking engine

#71

My heart goes out to this author, but you can tell even by his first table that he doesn't quite understand the mathematics of financial markets, the purpose of a hedge fund, how they grow etc. 1) It's plain by quickly looking at the allocation of capital in investment firms, that AUM is not made by performance; it's marketing. At best people invest when they believe a person is connected to inside information. Sayin…

> That markets are entropy machines rougher than a normal distribution, and any gains come directly from information. Isn’t this partially what this model is accomplishing with sentiment analysis? Also has there been a lot of investment into sentiment analysis for algo trading? I’m sure there have but references including books would be interesting.

No. Because to train the sentiment model you need estimable distributions.

Re: The anatomy of an ML-powered stock picking engine

#72
I'm curious what happens if you look at your returns and other metrics at different time scales, i.e. monthly and weekly, in addition to yearly. You can't make any argument based on a sample size of 2.

As someone who used to work in the industry, I am 99.99% confident that you cannot have any alpha with a system like this, you are basically flipping coins, as some other commenters have pointed out.

Re: The anatomy of an ML-powered stock picking engine

#73

Hi, fellow HN'ers! Author here, please let me know if you have any questions or thoughts!

Thanks for sharing! Really curious about ML stock market models as they seem extremely difficult to outperform the market consistently over time.

A few questions:

1) Were these stock picks for major stocks / ETFs? Or small market cap stocks?

2) How many people were subscribed to your newsletter?

3) What do you estimate the impact was of creating a “self-fulfilling prophecy” of entering a position and then recommending your subscribers take the same position?

4) Do you think your asset mix outperformed the market by picking high risk / high reward stocks in a bull market? Or picking safe stocks in a bear market? In other words, how do you know the engine wasn’t biased towards the market trend that happened to play out? For example, if I have a basket of tech stocks that would typically outperform in a bull market and flip a coin to buy or short them and guess right, I could outperform the market by chance. How did you account for this?

5) Have you backtested the engine to see what it would have returned in previous years? (Obviously on unseen data, rather than data it used for training)

Re: The anatomy of an ML-powered stock picking engine

#74
I know a bit about this industry and I have worked on some profitable systems. Honestly not a bad effort for someone working on their own with low-cost data. Don’t let the haters get you down. I would recommend you to pick up a more recent textbook on portfolio construction like Isichenko’s recent book.

Re: The anatomy of an ML-powered stock picking engine

#75
"steadily beating the S&P 500 for over a year on a weekly basis"

Can be achieved by chance alone.

If not chance, I can give you a strategy that would be highly likely to achieve such a result: it would take a lot of risk though!

I love it when people post this stuff to HN. Naive people try it, loose a bundle to market makers, then go back to their day job.

Re: The anatomy of an ML-powered stock picking engine

#76
post #61

Earlier quoted context omitted.

Some simply build a portfolio by copying those who can't be charged for violating market rules. Not sure why some folks find this strategy so controversial. =) Congress member holdings report: http://clerk.house.gov/public_disc/financial-search.aspx Senate member holdings report: https://efdsearch.senate.gov/search/

That's a very interesting idea. One obvious downside is a congressman might change their position faster than you can because they have advanced knowledge, so it would be riskier to hold the same position. Especially if that linked page updates even a little slowly. Secondly, they might have that position knowing exactly that - that they will have advanced knowledge and there it's worth having it at all. They know th…

This actually makes me think that an appropriate way to deal with congresspeople trading stock may be to require trade disclosure, say, a week in advance.

Re: The anatomy of an ML-powered stock picking engine

#77

Hi, fellow HN'ers! Author here, please let me know if you have any questions or thoughts!

EDGAR filings (structured text) is an area unto itself, I see you've limited yourself to quarterlies.

Across any market area (eg: mineral resources) there are thousands of documents released daily across multiple exchanges (via EDGAR, SEDAR, etc) ranging from two line advisories, to 4,000 page technical reports on projects | acquisitions, alongside the usual quarterly | yearly annual reports, etc.

There's plenty to do parsing common forms for generic changes (board members, board member share changes, etc) and market regime specifics (exploration property aquisition) and trends (series of related aquisitions) for those that like the weeds.

Some might argue that 'understanding' these patterns lead the changes in stock price movements, and give insight wrt weathering short term changes for longer term returns.

Re: The anatomy of an ML-powered stock picking engine

#78
post #74

I know a bit about this industry and I have worked on some profitable systems. Honestly not a bad effort for someone working on their own with low-cost data. Don’t let the haters get you down. I would recommend you to pick up a more recent textbook on portfolio construction like Isichenko’s recent book.

Thank you for the note! Just picked up Isichenko from the online bookstore we all love-hate.

I'd love to get in touch (as per your HN profile) - my email is am(at)principiamundi.com

Re: The anatomy of an ML-powered stock picking engine

#79
This was a very enjoyable read. I built a nearly (architecturally) identical system a few years back that also had to be scrapped for different reasons. This brought back a lot of memories. The sanity checks, the index reconstitution issues, dealing with the insanity of security identification and tracking through time.

The fun cases are the ones where it's not even clear what the right answer truly is, e.g. company A spins out company B, and then 5 years later they re-merge. Who's time series and associated data is "the" canonical one? The data vendors often try to give their answers to this question, but maybe their answers don't make sense for your analysis.

Then there's the fact that a lot of vendors don't really do point in time correctly. They like to go back and helpfully revise data points for you that they or the company initially misreported. This is all well and good except that if you were trading for real, you wouldn't have known the correct information at the time, and so any backtest based on the updated information will be invalid. Vendors are a bit better now about providing true point in time data sets, or at the very least accurately describing when they are/aren't doing this. But we had a few cases where they said they were, but they definitely weren't.

Re: The anatomy of an ML-powered stock picking engine

#80
post #75

"steadily beating the S&P 500 for over a year on a weekly basis" Can be achieved by chance alone. If not chance, I can give you a strategy that would be highly likely to achieve such a result: it would take a lot of risk though! I love it when people post this stuff to HN. Naive people try it, loose a bundle to market makers, then go back to their day job.

> "steadily beating the S&P 500 for over a year on a weekly basis"

If you're going to make a claim like that, you should actually follow up with the calculations. When you do that, you'll realize that the issue is quite a bit more complex than this shallow dismissal.

He has very low correlation to the index, which means he's not just levering beta and getting lucky on a trending market. His standard deviation is smaller than the index, which means he didn't just make one large and lucky bet. The evidence that he has real alpha is certainly not incontrovertible, but the numbers look quite good.

It also doesn't appear that he cherry picked his reporting/aggregation cadence, because he sent out a weekly newsletter ex ante, and all his stats are reported weekly. He could still just be lucky, but his numbers are much better than this sort of dismissal would imply.

One real risk is that, in some implicit way, he's pursing a negatively-skewed strategy. That is, one that has a latent large downside risk. Strategies like this can produce very good looking numbers for longish periods, but still have ultimately negative alpha. Judging whether or not that is the case here is hard without more detail, but nothing he says in the writeup indicates to me that that is the case here.

Post reply on HN