I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried:
a) "Berlin":
1. The movie festival "Berlinale"
2. The Wikipedia entry about Berlin
3. Something about a venue "Little Berlin", but the link resolves to an online gaming site from Singapure
4. "Visit Berlin", the official tourism site of Berlin
5. The hash tag "#Berlin" on Twitter
6. "1011 Now" a local news site for Lincoln, Nebraska
7. "Freie Universität Berlin"
8. Some random "Berlin" videos on Youtube
9. The Berlin Declaration of the Open Access Initiative
10. Some random "Berlin" entries on IMDb
11. A "Berlin" Nightclub from Chicago
12. Some random "Berlin" books on Amazon
13. The town of Berlin, Maryland
14. Some random "Berlin" entries on Facebook
15. The BMW Berlin Marathon
b) "philosophy"
1. The Wikipedia entry about philosophy
2. "Skin Care, Fragrances, and Bath & Body Gifts" from philosophy.com
3. "Unconditional Love Shampoo, Bath & Shower Gel" from philosophy.com
4. Definition of Philosophy at Dictionary.com
5. The Stanford Encyclopedia of Philosophy
6. PhilPapers, an index and bibliography of philosophy
7. The University of Science and Philosophy, a rather insignificant institution that happens to use the domain philosophy.org
8. "What Can I Do With This Major?" section about philosophy
9. Pages on "philosophy" from "Psychology Today". I looked at the first and found it to be too short and eclectic to be useful.
10. The Department of philosophy of Tufts University
c) "history"
1. Some random pages from history.com
2. "Watch Full Episodes of Your Favorite Shows" from history.com
3. Some random pages from history.org
4. "Battle of Bunker Hill begins" from history.com
5. Some random "History" pages from bbc.co.uk
6. Some random pages from historyplace.com
7. The hash tag "#history" on Twitter
8. The Missouri Historical Society (mohistory.com)
9. Some random pages from History Channel
10. Some random pages from the U.S. Census Bureau (www.census.gov/history/)
d) "Caesar"
1. The Wikipedia entry about Caesar
2. Little Caesars Pizza
3. "CAESAR", a source for body measurement data. But the link is dead and resolves to SAE International, a professional association for engineering
4. The Caesar Stiftung, a neuroethology institute
5. Some random "Caesar" books on Amazon
6. Hotels and Casinos of a Caesars group
7. A very short bio of Julius Ceasar on livius.org
8. Texts on and from Caesar provided by a University of Chicago scholar
9. (Extremely short) articles related to Caesar from britannica.com
10. "Syria: Stories Behind Photos of Killed Detainees | Human Rights Watch". The photos were by an organization called the Caesar Files Group
So what I can see are some high ranked false positives that are somehow using the search term, but not in its basic meaning (a3, a11, b2, b3, d2, d3, d4, d6) or not even that (a6). Some results are ranking prominently although they are of minor importance for the (general) search term (a9, a13, b7, b8 -- perhaps a15 and d10). Then there are the links to the usual suspects such as Wikipedia, Twitter, Amazon, etc. (a2, a5, a8, a10, a12, a14, b7, c5, d1, d5); I understand that Wikipedia articles are featuring prominently, but for the others I would rather go directly to eg. Amazon when I am interested in finding a book (or use a search term like "Caesar amazon" or "Caesar books"). Well, and then there are the search results that are not completely off, but either contain almost no information, at least compared to the corresponding Wikipedia article and its summary (b4, b9, d7, d9), or that are too specific for the general search term (c1, c2, c3, c4, c6, c9, c10).
That leaves me with the following more or less high quality results (outside of the Wikipedia pages): a1, a4, a7, b5, b6, b10, and d8. The a15 and d10 results I could tolerate if there had been more high quality results in front of them; but as a fourth and second, respectively, good result they seem to me to be too prominent. Also in the case of "Berlin" a4 should have been more prominent than a1, and a7 is somewhat arbitrary, because Humbolt University and the Technical University of Berlin are likewise important; what is completely missing is the official Website of the city of Berlin (English version at www.berlin.de/en/).
All in all, I would say that your ranking algorithm lacks semantic context. It seems the prominence of an entry is mainly determined by either just being from the big players like Twitter, Youtube, Amazon, Facebook, etc. or by the search term appearing in the domain name or the path of the resource, regardless of the quality of the content.