Live data from Hacker News

Four Areas of Legal Ripe for Disruption by Smart Startups

lawtechnologytoday.org

51–60 of 72 posts

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#52
post #32

I'd be interested in seeing an IDE for contract drafting. Or something like an Excel mapper that can show me how all the provisions and definitions in a contract are interrelated.

Definitely interesting, and in a few cases I have built this internally for myself before and it generally is useful (key phrase) when it can be built. I did this primarily to keep track of (1) defined terms (each place a defined term is used, and when a defined term includes another defined term), (2) section references (where any section is referred to elsewhere in the document so that if that section changes, what other things need to be impacted) and (3) payment mechanics (a spreadsheet that uses the defined terms of payment mechanics and allows you to plug in real numbers to see how the money flows). This is helpful to understand a specific concept or mechanic.

Usually though, the moment you move into a transaction of even medium complexity, while this might be helpful it can't be ultimately relied on - if you were to track a defined term or a section reference, a change to that term or section reference would impact not only the places where that term or section reference is specifically used, but also where a concept depends on such term or section reference. A basic example of the change of an actor from singular to plural (originally there was one purchaser, now there are two purchasers) - then, every pronoun and verb would need to be changed to plural form. This is why you sometimes find contracts wonkily sticking with a plural defined term when really there is just one entity/person described by the term, or vice versa.

This mostly points to what I always say about the challenges to true legal disruption: common law, statutes and contracts all depend on language subject to interpretation (and in the case of common law, interpretation IS the entire name of the game - see every supreme court case of the past 30 years to see how much interpretation varies), and unless I am missing some major breakthrough, we have not come to a point where language is understood systematically enough to truly "hack" complex legal concepts as currently drafted (aka. using language).

I could, however, imagine a new legal system that depended entirely on data, numbers and systems instead of historical language, but it would first require us to all agree that the system would govern and agree on the rules (or lack thereof) of interpretation, which given the vested interests most of society has in the current system, would be pretty challenging to implement.

That being said, you can see how this spectrum works by comparing the common law system (like US/UK, which relies heavily on interpretation of judicial opinions) and the civil law system (France, Korea, etc., which relies heavily on a more formulaic interpretation of statutes) - way more costly litigation and more politics in common law systems when compared to civil law systems. The trade off is that common law is (arguably) more dynamic (judges can overturn statutes unilaterally - civil rights, etc.), whereas civil law systems require legislatures to make changes to statutes.

Full disclosure: I used to be a lawyer, so I am biased.

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#53
post #6

This paralegal I know told me 2 years ago "I wish someone would automate discovery because it sucks right now". I wish I knew absolutely anything about it.

This is called eDiscovery, and people have been working on it for a good while, certainly longer than two years (I worked at a startup in this space, 2009-2010).

It's interesting to note the first book on eDiscovery came out in 2004. Didn't sell many copies as you can imagine, and was pretty ahead of its time.

Pretty shocking when you think the first book on digital evidence didn't come out until 2004. You want a legal field ripe for disruption? It's definitely the eDiscovery space.

Here's the book I was referring to:

http://legalsolutions.thomsonreuters.com/law-products/Treati...

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#54
post #49

I currently write software for an e-discovery company. Most tasks that our software is expected to be able to perform are simple-sounding tasks, at first glance, such as ... 1. extracting documents from within other documents (attachments out of an email, files out of a zip, embedded excels out of a word doc, images out of a powerpoint, etc) 2. convert all said documents to some kind of standard media format so that…

Feature requests in e-discovery bloat your original software out of all proportion. Nevermind getting past the original part of effectively searching large troves of data in different formats.

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#55
post #2

It's hard to get excited about software for lawyers, and I think that's why Disco has flown under the radar a bit, but I think these guys are going to be huge. They've made exponential improvements in e-discovery software.

Having worked in e-discovery for many years, "exponential improvements" is a huge overstatement. What they have is in pretty much any e-discovery product on the market. And they are missing a huge piece--predictive coding and advanced analytics (email threading and near-dup are EXTREMELY common). If you are in NLP, ML and/or IR, the legal industry is probably one of the most exciting places to be. Huge datasets, avai…

Interesting. What type of advanced analytics do you mean?

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#56

I am quite interested to know more about Judicata. It was cofounded by Blake Masters (coauthor of Zero to One). Anyone have any info on it?

As somebody doing a very similar thing to Judicata elsewhere in the world, I have looked into Judicata quite a bit. Although they are in stealth mode, there is some information out there and I have been able to infer a bit about what they are doing. Essentially, they are seeking to use NLP to retrieve certain information from legal texts. The primary piece of information they seek to retrieve is the legal claim, and certain elements surrounding it (e.g. 'breach of s X of Y statute'). They also seek to retrieve a number of other bits of information from case law using NLP.

In line with Peter Thiel / Palantir's philosophy that the human brain is an amazing machine not to be supplanted by computers, but one that should be used to its fullest, augmented by computers, Judicata's software involves using NLP as much as possible, then feeding or 'striating' that information to lawyers or legally trained people, depending on the complexity of the information extracted, for their confirmation. This is in any case necessary because NLP cannot get close to 100% accuracy for the information they are trying to extract, and you need 100% accuracy in the legal domain (e.g. it would be unacceptable to get the legal claim wrong, c.f. Google search).

One consequence of structured legal texts is improved search. What many don't realise is the degree with which structured search on legal texts will improve legal research. E.g., if I want to find all cases in the last 10 years where the plaintiff claimed breach of duty in an occupiers' liability suit, I simply cannot. To find that batch of cases (accurately) would take me hours. If the legal claim was a structured piece of information, I could just search for it. As an ex-lawyer and ex-legal researcher, the number of hours that could be saved per lawyer per year could easily be in the hundreds, and this is at charge out rates of $300-$1k per hour. This is, similarly to the above comments, in line with Thiel's investment thesis to 'improve something 10-fold' or 'make a quantum advance to cause adoption / change consumer behavior'. I think most people seriously underestimate how significant of an improvement structured search would be.

The other thing that Judicata are flying under the radar about, a little bit, is the ability to use structured legal information for other purposes. High on the list is analytics, which Itai Gurari mentioned at the end of a talk, but merely in passing as if it was inconsequential. I think this is pretty clearly a multi-billion dollar market waiting to be made. If you look at what similar firms are doing in niche areas of law, e.g. Lex Machina, and look at what they are charging, and extrapolate the types of questions you can answer with structured legal information, the potential becomes clear. Again, this is in line with Thiel's investment thesis to 'create a market a dominate it, rather than compete in an existing one'.

The primary difficulty for Judicata or somebody undertaking to do the same thing is that the task is mammoth in just about every respect. As such the optimum strategy is likely to attack a niche jurisdiction and then build out the product. You can't go 'full-lean', because you need at least a semi-complete data set, but you can start 'small'. Hence, Judicata have been working on a niche jurisdiction of law as their first project: Californian Employment Law. While I am not in that jurisdiction (not even in America), that seems to me to be a very reasonable area of law to start with given that most legal claims (I think) are found in California's employment law statute (as opposed to other areas where the legal claims are found in Judge-made common law). Furthermore, there are a ton of neat pieces of information in employment law which you can structure, e.g. in a discrimination case, what factor was the plaintiff allegedly discriminated on - race, age, gender, etc. Finally, uptake would be high among employment lawyers who research at reasonably frequent intervals and have a practical need for more accurate search; compare this to constitutional law for example.

While part of the reason they are operating in stealth mode is simply because it takes so long to build up a semi-complete data set, I think the other part of the reason is because the biggest risk for such a firm is that Lexis, West or Bloomberg will start doing something similar. Imho, it's likely they will eventually but the risk of Judicata catalyzing that process is pretty small.

There are a few other firms operating in this space but with fundamentally different philosophies. My view is that these other firms are simply taking the wrong approach and simply want to release a product and build on it now in the lean tradition. Judicata's product is the type of product where the question is not whether there will be adoption, but rather whether or not you can actually build the product on your budget and in the time frame required. Imho, if Judicata can successfully create what they are planning on, it will flatten their competition. The real question is whether they can.

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#57
post #11

The article mentions some interesting things, but I have a few quibbles: > But even with today’s modern communication tools, both customer experience and lawyer workflow have remained stagnant. At a large firm, legal practice is unrecognizable compared to even 10-15 years ago. Everything is electronic: filing and docketing, document collection/scanning/OCR, legal research, document management (DMS + version control).…

> PageRank, for example, works great when everyone searching for "skiing near Tahoe" is looking for the same popular pages. But when you're doing legal research, a lower-court case that directly addresses your issue but isn't widely cited is much more valuable than a highly-cited Supreme Court case that doesn't address your issue.

This is true where you are doing fine-grained research, but there are plenty of lawyers who are constantly delving into new areas of law that are adjacent to or tangential to their primary area of interest or practice. When this happens, it can be phenomenally useful to get up to speed on the area of law by seeing which decisions have the highest PageRank, or PageRank weighted by certain factors, etc.

In this way, one way to think about the startup that implements a PageRank algorithm is that it is not disrupting electronic legal research, but rather disrupting legal textbooks. In my experience, the fastest way to get abreast of a new field is to find the leading textbook. One could imagine a sophisticated database which could altogether remove the need to consult a textbook by painting a picture of the leading cases, key pieces of legislation, which sections are most often referred by which cases in which context, etc.

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#58
eDiscovery is not ripe for disruption. That ship has sailed.

eDiscovery has largely been solved for most corporate environments. There are tools to collect data in a defensible manner, to "process" (i.e., index) it, and to review it. There are even some products that aggregate these functions together, however, it must be well-noted that each of these functions has a different user/customer and occurs at a different timeframe in the discovery process.

Many of the dominant tools do have their warts. But the money that was once in this space--the eDiscovery collection product I wrote sold for a couple million to its first customer--is no longer there. Prices have dropped dramatically and its now a commoditized market. So you'd have to work very hard for very little gain to displace any of the dominant players.

Note that TFA was written by investors in a new eDiscovery startup and TFA seems mostly like latent marketing for them. I don't know anything about them--good luck and all that--but I'm very familiar with the space and I don't envy them.

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#59
post #49

I currently write software for an e-discovery company. Most tasks that our software is expected to be able to perform are simple-sounding tasks, at first glance, such as ... 1. extracting documents from within other documents (attachments out of an email, files out of a zip, embedded excels out of a word doc, images out of a powerpoint, etc) 2. convert all said documents to some kind of standard media format so that…

The other thing is that you spend 10+X the initial effort of the feature for error handling. I sometimes envy all these "Internet" programmers who only have to deal with the web and don't have to worry about esoterica like, e.g., \x80 being the space character in old WordPerfect files.

Re: Four Areas of Legal Ripe for Disruption by Smart Startups

#60
post #6

This paralegal I know told me 2 years ago "I wish someone would automate discovery because it sucks right now". I wish I knew absolutely anything about it.

I know some about it. As it stands right now, law firms need to find, trust, and pay forensic investigators. The firm has to leave some data collection up to the investigator, and then have some data turned over. By data, I mean hard drives, images of hard drives, images of network shares, email exchanges, etc. The law firm has to pay for the data collection, the disk space to store the data, the transmission of data…

I'm a computer forensics and eDiscovery guy. The forensics part of eDiscovery is really overblown. It's the bogeyman, and not much more.
Post reply on HN