Live data from Hacker News

The anatomy of an ML-powered stock picking engine

principiamundi.com

51–60 of 109 posts

Re: The anatomy of an ML-powered stock picking engine

#51
post #50

Someone asked about how difficult it is to get outside investment.... It's usually very difficult and it takes a lot of money to run a proper fund. Let's say you raise $50M. You can maybe charge 1 and 20,meaning you get 1% of assets each year for running the fund and 20% of profits. 1% of $50M( and keep in mind this is a large raise for someone without a track record on the sell side or inside another fund) give you…

This is actually way too optimistic. Your first 1-2 seed investors will: - Only pay 1 and 10 (1% fixed fee and 10% of PNL) - They will also get ownership of the actual fund management firm and will get that in the form of 20% of REVENUE (not equity, revenue, think about that) This is one reason new fund formation is way down. The economics are bad for years. Know a bunch of HF people that started vc-backed tech firms…

BTW, data costs also too low.

Just a BB terminal around 30k and a lot of extra data from BB costs extra (can be 200-300k per additional product).

For quant strategy probably looking at 500k up to 2M for data initially. And you will likely be at a disadvantage to existing firms that have been collecting data for years.

And that is at the low end. Spent many millions per year for 1 strategy at last large firm. And that was small fraction of total firm spend.

Re: The anatomy of an ML-powered stock picking engine

#53

Earlier quoted context omitted.

Thank you for appreciating the article; I tried to disclose all that I could! 1. Yes, I did put my own money in it (low 6 figures). 2. It went as described in the article - for the capital I allocated to Didact, I beat the market (SPY) by ~20% since inception. 3. If I understand your question correctly, this would be the equivalent of the payoff on an optimal lookback option ( https://en.wikipedia.org/wiki/Lookback_o…

>2. It went as described in the article - for the capital I allocated to Didact, I beat the market (SPY) by ~20% since inception. This seems extremely hard to believe. You should be running a multi-billion $ Quant fund if this is the case. The idea that you would try to push this as a newsletter rather than just taking investor money and becoming a billionaire literally makes the story seem farcical.

It is very easy to believe.

I could have flipped a coin, gone long or short at beginning of this year.

I would have had a 50% chance of outperforming the market by 40% this year (given it is down roughly 20%).

Re: The anatomy of an ML-powered stock picking engine

#55
My heart goes out to this author, but you can tell even by his first table that he doesn't quite understand the mathematics of financial markets, the purpose of a hedge fund, how they grow etc.

1) It's plain by quickly looking at the allocation of capital in investment firms, that AUM is not made by performance; it's marketing. At best people invest when they believe a person is connected to inside information. Saying you have an ML advisor is really just a pre-req to these people.

2) Is that allocation stupid? No, it's not, because actually the powers of mathematics and by extension ML are intrinsically limited for investment returns because they are fat-tailed . For example this author quotes a realistic sharpe (0.8), but didn't calculate the standard deviation in his sharpe, which I would bet a large sum was _at least_ 0.8. Ie: he doesn't really know what his sharpe is. This is because equity assets behave like a student-t distributions with a degree-of-freedom parameter ~2 or less . Ie: higher moments such as uncertainty in sharpe, literally do not exist or converge and are unknowable. The only exception is if your strategy explicitly cuts off tails.

Once you understand 2) you begin to understand that there's no such thing as a real quant fund (ie a fund which truly makes money predictably using models) which doesn't trade a liquidity limited book that has quite advanced hedging. Wealthy people are aware of this, which is why the author can't market this product.

If you're doing something silly like holding equities without tail risk control, you literally cannot be quantitatively investing. You are just slowly rediscovering what Kelly, Bergomi, Mandlebrot, Bernay's etc. realized with a little deep thought over pen and paper (while clumsily writing boilerplate software.) That markets are entropy machines rougher than a normal distribution, and any gains come directly from information. (see: Kelly: "a novel interpretation of the information rate".)

For a high latency (ms) market data feed, the returns on information are very very small. Markets are efficient.

Re: The anatomy of an ML-powered stock picking engine

#56
post #16

this is very cool! where did you get your data from and how's the transition to airflow?

There are commercial feeds available via Nasdaq DataLink (FKA Quandl). I also bought bulk historical data to feed through my backtester (I haven't talked about this in the post; it was getting to be a bit too long).

Let's get a write up of your backtesting framework too please! Terrific post @muggermuch - thank you!

Re: The anatomy of an ML-powered stock picking engine

#57

Earlier quoted context omitted.

There are commercial feeds available via Nasdaq DataLink (FKA Quandl). I also bought bulk historical data to feed through my backtester (I haven't talked about this in the post; it was getting to be a bit too long).

Let's get a write up of your backtesting framework too please! Terrific post @muggermuch - thank you!

:) Will do! Thanks for the encouragement!

Re: The anatomy of an ML-powered stock picking engine

#58
What's the market beta? What's the average turnover/holding period? How are transaction costs modelled? What features explain most of the variance? How are they related to known factors? What's the beta hedged performance?

These are all things I'd want to know before deploying something like this. (Perhaps some mentioned in the post, might have missed them.)

To first order, I'd forget about fat tails and similar popular concerns. They matter, but not as much as structurally understanding what this model is up to. Perhaps one feature is explicitly selling tails? That might answer it already.

Re: The anatomy of an ML-powered stock picking engine

#59
post #11

Earlier quoted context omitted.

Great question. If I beat the market by 20% (say SPY generated 0% for the year, very optimistic at this point), and I have allocated $100k to this, I make $20k before taxes. That's less than minimum wage. Meanwhile, allocators expect a track record of at least 3-5 years. Ideally, if I have an asset, I'd like to extract as much revenue as I can. Hope this makes sense.

How difficult is it to get investors when you can show your model beats the market consistently? Of course, they have to check your not trading a strategy with extreme tail risk, but here it sounds like that's not the case?

Very, because historically a levered long position in the market beats the market consistently over long periods. Beating the market turns out not to be a very interesting metric to sophisticated investors.

Re: The anatomy of an ML-powered stock picking engine

#60

Every gambler thinks they have a system, but often fails to recognize a game is unfair long before they arrived. lol =)

You can think outside the box to beat the unfair game but then you end up in jail.

Some simply build a portfolio by copying those who can't be charged for violating market rules. Not sure why some folks find this strategy so controversial. =)

Congress member holdings report:

http://clerk.house.gov/public_disc/financial-search.aspx

Senate member holdings report:

https://efdsearch.senate.gov/search/

Post reply on HN