I've started to re-read Designing Data-Intensive Applications (this year they published a second edition of this classic book). I like to call it "the wild boar's book."

Here are some comments on chapter one, "Trade-Offs in Data System Architecture."

  • How do we reason about trade-offs?
  • Does AI help us consider different scenarios?

Different people need to do very different things with data. Is AI another one that needs a specific type of data? Let's say applications need transactional/operational data systems. Business people need analytical data systems. OK... LLMs need another kind of approach? Maybe 'data silos' isn't such a bad term here.

"Consequently, organizations face a need to make data available in a form that is suitable for use by data scientists..."

From data warehouses to data lakes because of data scientists' needs? Weird.

  • How do new technologies emerge?
  • What are the conditions that allow a new technology to emerge, and what are the conditions that allow it to be reproduced?

We changed from capacity planning and performance optimization to financial planning and cost optimization. The cloud changed the roles; the AI boom is doing the same. - What could we learn?

Recognition of the effects (not just side effects) that computer systems have on people and society.

We store data because we think that its value is greater than the costs of storing it. The risks involved push us toward data minimization. Is focusing on risks the only way?