Live data from Hacker News

Ask HN: How to understand the large codebase of an open-source project?

news.ycombinator.com

11–20 of 49 posts

Re: Ask HN: How to understand the large codebase of an open-source project?

#14
Here is what I don't do:

1. Read some code

2. Try to understand how it works

3. Repeat

Here is what I do:

1. Try to figure out what the code might be in advance using the information you have. (For example: I know nothing but the fact that it's a spreadsheet. Then figure out in your mind how the basics of a spreadsheet might work.)

2. Now read a little bit of the code. Compare with what you were thinking. If it matches, go to 3. If it doesnt match, figure out why by reading the code and by thinking more.

3. Repeat

Note the two processes are relatively similar because step 2 of the former process is a little bit like step 1 of the later process. Just try to focus on figuring out first, read second. Figure out first, read second. It's an active approach which makes you work more, and the more you work, the faster you go - or some benefits of that sort.

I actually wonder if people do that.

Re: Ask HN: How to understand the large codebase of an open-source project?

#16
You can use the debugger on low level api calls to get pretty much anywhere in the codebase. If you want to find whats changing a label to "foo" you can hook into every set_Text call and put a conditional breakpoint on all label changes to break on "foo", then just go up the callstack to find the logic. This strategy works on network interfaces and file interfaces as well. I abused this on our 2M+ SLOC legacy codebase and it has saved me many hours.

Also use version control to identify the most commonly edited files in the project. These are usually the files that are doing all the work (80/20 rule) and you likely need to know of them.

git log --pretty=format: --name-only | sort | uniq -c | sort -rg | head -10

Re: Ask HN: How to understand the large codebase of an open-source project?

#18
Assuming you know how to use the product: write down a path within the software that's intuitively familiar to you. Then follow that same path in the code, starting from main() or equivalent.

When you trace your own usage footsteps like this, it's often amazing how much goes on behind the scenes that you never realized.

Re: Ask HN: How to understand the large codebase of an open-source project?

#19
Off the top of my head:

- Count the lines of code with find | wc, get a sense for what's there, and what language it's written in. The biggest file in the project is usually worth a look -- it is often where the "meat" is. Read the function names.

- Use the program. grep for strings that appear in the UI in the source code. That's a good place to start reading. Read function names.

- strace the program. What system calls does it make when? ltrace is also sometimes useful, although it also gives a ton of output.

- Look at header files. Understanding data structures is often easier than understanding code.

- Look at commit logs. Those are hidden "comments". And reading diffs can be easier than reading code.

- Do a "log" or "blame" on the file. How has it evolved?

- Start reading main(). This often reveals something about the structure of the program. Even just finding main() in many programs is a good exercise :) Sometimes it's a little hard to find.

- Make sure to build it. And if you can, look at the build system. How is it put together? Most build systems are pretty darn unreadable. I don't really know how to read autoconf, and GNU make is tough too. Forget about cmake :) But sometimes this can help.

I haven't gotten that far with this, but I tried uftrace recently and like it:

https://github.com/namhyung/uftrace

You can think of it like a dtrace that knows about every function in a C or C++ program.

-----

I want to try some kind of code explorer thing. I saw this in a CppCon video and on HN:

https://www.sourcetrail.com/

And older ones like:

https://www.sourceinsight.com/

But somehow I get by with Unix tools. I think this is because I feel like building the project in a way to accomodate the source browsers might be a big pain.

Counterpoint: I think the hardest part of understanding a project is usually the build system :-) I don't have too much of a problem with reading C, C++, Python, or (sometimes) JS code. Volume is always a problem, but I can read a specific function pretty easily. But the build system is where things get ugly, in my experience.

Also, reading multi-threaded code requires some special consideration. grepping for every place that threads are started is a good idea.

Post reply on HN