Live data from Hacker News

U.S. universities, rich in data, struggle to capture its value, study finds

newsroom.ucla.edu

31–40 of 97 posts

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#31
post #19

Universities already waste vast sums of money of vice-provosts, middle management, administration, etc. There's no need to encourage them to start treating their mission as one of data mining in order to "capture" more value from the students.

While I agree with your general skepticism, I also believe there are ways these data could benefit the students. Extracting value from operational data could mean modernizing curriculum, offering more office hours, etc.

I think most of the value of "data" is captured in empowering individual contributors to observe their working conditions and the impact of their actions and adapt their strategy in response. The power of the central office spreadsheet wielders is very secondary. (Which is of course not to say that spreadsheets are not valuable, as any teacher knows.)

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#32
post #25

Obviously I’m 20 years past university and nearer being the paying parent, but why oh why should university “capture the value of data”. They should offer affordable education (if you believe in the value of tertiary education) and affordable signaling of abilities (if you don’t). A recent post by John Cochrane [1] somewhat pointed out the absurdity in US education, Stanford specifically. Stanford itself mentions 157…

Public schools depend on a lot of administrative support outside the physical school its self. Buss drivers for example aren’t managed at the individual school level because they transport students to multiple different schools.

This extends through a huge range of administrative functions for everything from calling snow days to collecting taxes to pay for the school etc.

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#33

I can't access the original study however taking data, explicitly naming security cameras for example, to "use it or merge it with data from external parties such as publishers or public or private sector organizations" will surely not seriously degrade student and staff privacy, right? The rush to "exploit" data reminds me of the dot com hype. It's one thing to use available data to make more informed decisions abou…

> degrade student and staff privacy

That's the primary value to be captured right there! Their privacy is highly valued, monetarily:

> authors contend that universities have been slower than organizations in other _economic_ sectors

Education in the US is just pure business.

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#34

As a data scientist, I think most data is useless, but there is an addictive, video game-like quality to throwing lifeless spreadsheets into a machine and having colorful visualizations come out. It's kind of like a very boring video game for adults that makes them feel like they're working, when they're actually just enjoying colorful abstract shapes and colors. To be honest, this is probably a sizable piece of why…

Agreed! It all probably started with big tech promoting the “data-driven decision” paradigm. Of course, in many cases this approach is effective, but it’s not a panacea and has its limits. It’s tempting to interpret availability of any data as an amazing untapped resource, but in many cases analyzing it is just a massive waste of resources and could be more effectively replaced with traditional tools (surveys and such).

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#35

There's a troubling tone behind this article. It's a prophecy that demands to be fulfilled. Starting with the premise "Data is useful", it proceeds to pick at all the ways we've failed to make it useful... and we're damned well going to make it useful if it kills us! Maybe, just maybe (for those that dabble in the sceptical, explanatory game we call science) it might be that "bare data" has little use within certain…

I think it's simpler: imagine you're the NSA and decide to transcribe every phone call ever made to text: you'll spend gazillion in storage, connectivity and processing for the capture, and spend a gazillion squared on post-capture text analysis to find "something", like "who is communist in Atlanta", and not even be able to make it find the communists pre-emptively, before it's too late and they already donated $3 to a local chapter.

There can be too much data. There are questions you cannot answer even by collecting all data, unless you have no time or food cost constraints but the value of your answer will decrease with the time distance from the moment the question was asked. In a million year you might know for sure who was communist in Atlanta during the 2000s, year by year, block by block, but you may care less by then.

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#36

Earlier quoted context omitted.

While I agree with your general skepticism, I also believe there are ways these data could benefit the students. Extracting value from operational data could mean modernizing curriculum, offering more office hours, etc.

I think most of the value of "data" is captured in empowering individual contributors to observe their working conditions and the impact of their actions and adapt their strategy in response. The power of the central office spreadsheet wielders is very secondary. (Which is of course not to say that spreadsheets are not valuable, as any teacher knows.)

In an ideal world - yes. But seeing how overworked professors and other staff members already are in top universities, I doubt this is top of mind for them. There are courses with hundreds of enrolled students, weekly homework, practical projects, etc. IMO somewhat centralized DS tools are more suited to handle this load than an individual contributor (a professor in this context).

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#37
Working in a top-tier computer science department, I find our ability to answer basic questions about the health of our degree program fairly troubling. I think non-academics may be surprised by how much we don't know, and how little useful and continuous data analysis is taking place.

For example: What is our retention rate? Meaning, what percentage of students who start our degree programs complete it. A fairly standard and important indicator of program health. Next, break this down by various cohorts: What is our retention rate among women? And so on. Heck, frequently we can't even answer questions about the current gender ratio within our program—and this is something that has been a focus of our diversity efforts recently.

I've had people say with a straight face that we _cannot_ calculate retention because we don't know when students leave our program. But of course someone knows this! And I've been able to produce rough estimates even given the limited data that I have access to. But a lot of educational data is fairly siloed, and frequently the people assigned to perform these tasks don't have much training and tend to give up quickly.

I suspect that many departments just don't have anyone assigned to do even basic educational data analysis on a regular basis, and with access to enough data to run interesting reports. My department is in the process of creating a faculty leadership role around academic data analytics, but my sense is that this will be a very unusual position. (And don't worry—it'll be filled by a faculty member, and not a new administrator.)

And don't even get me started about student evaluations of teaching. Yes, we give a survey at the end of every semester and ask students whether they liked a particular course and professor. No, those answers have very little to do with how much they actually learned. Yes, we could measure learning in other better ways—success in downstream courses, for example. No, people don't tend to do that.

There's a lot of room for improvement here, just working with the data we already have. No need for additional "telemetric signals".

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#38

This article feels like it took hundreds of words to say nothing at all. The funniest part was the example of security cameras. What big breakthroughs are universities hoping to achieve with this data?

Security cameras in universities seem to be accomplishing… nothing. I am getting very frequent emails (2-3 per week) from one of the top US universities with blurry unusable screenshots of perpetrators stealing items, breaking in, assaulting students, etc. Makes me think - should they not invest into better security instead?

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#39
post #12

I have no doubt the problem is worse in academia, but to be quite honest nobody is really doing a good job at curating, governing and unifying data across large organizations (at least without extreme expense.) Just look at the treadmill of architectures designed to "solve" this problem: integrated databases to "data warehouses" to "linked data" to "data lakes" to "data fabric" to "data mesh." Organizations that succ…

Observation/mini rant: unfortunately the data industry is both very susceptible to fads and hype, and does not have widespread standardization or a generally-accepted set of best practices. The subset of data "influencers" whose discourse occurs mostly through Twitter and Substack have especially heavy sway in what is the current Data Big Thing. These people introduce new buzzwords and ride them to the mainstream, or redefine what were previously more-or-less anchored down concepts into new interpretations. It feels tepid and arbitrary, almost postmodern. On top of this, much of the "thought leadership" is being driven by individuals who lean heavily towards the soft-skill side. So we have a heavy overindexing on strong opinions around organizational methodologies, team structure and roles/responsibilities, and other Big Ideas without much engineering representation or consideration. What spawns from this are things like the "data mesh" and the bastardization of both data concepts and clear communication/thinking. The data mesh white paper perfectly encapsulates this[1]. I challenge anybody to try to read it and understand what the hell the author is even vaguely trying to say after a few passes.

[1]: https://martinfowler.com/articles/data-mesh-principles.html

Re: U.S. universities, rich in data, struggle to capture its value, study finds

#40

As a data scientist, I think most data is useless, but there is an addictive, video game-like quality to throwing lifeless spreadsheets into a machine and having colorful visualizations come out. It's kind of like a very boring video game for adults that makes them feel like they're working, when they're actually just enjoying colorful abstract shapes and colors. To be honest, this is probably a sizable piece of why…

> It's kind of like a very boring video game for adults that makes them feel like they're working,

Absolutely love it.

Let's distinguish a few things though. "Data science" seems like a pretty weird name. I mean, it's just "Science" right. Of course there's statistics, mathematics, signal processing, systems analysis, machine learning... all the good things that you and I are into.

But how does this get huddled uncomfortably beneath the umbrella "Data science"?

I think the answer is found by asking about the ends of data science, the old Cui Bono?

There's the raw entertainment value you mention. It's cool to have knowledge and visualise it. Sensors, transducers, processing is fun.

Then there's legibility. That is political and is about control.

What most scientists are doing with data is either hypothesis testing or combing for causal relations to then abductively feed back into hypothesis generation.

What most business people are trying to do is optimise, and adjust constraints and parameters. It's modelling for the most-part. It's ancient and goes back to linear analysis and regression from before the last century.

Security people are looking for stress signifiers, suspicious patterns with various triggers, selectors and tripwires.

Financial people want fortune telling. They want the models to extrapolate into beautiful hockey sticks.

Within any organisation we may need to do one, a few, many or none at all of the above. The problem then is that "Valuable data" is such a broad, open prospect it seduces gushing, credulous administrators into valuing the process, and the tools, but not the ends.

Post reply on HN