Giulia Lanzafame

Giulia Lanzafame

5 posts

Implement an enterprise-ready data lakehouse architecture with Spark and Kyuubi

Here at Canonical we are excited to announce that we have shipped the first release of our solution for enterprise-ready data lakehouses, built on the...

Accelerating data science with Apache Spark and GPUs

Apache Spark has always been very well known for distributing computation among multiple nodes using the assistance of partitions, and CPU cores have always...

Apache Spark security: start with a solid foundation

Everyone agrees security matters – yet when it comes to big data analytics with Apache Spark, it’s not just another checkbox. Spark’s open source Java...

Spark or Hadoop: the best choice for big data teams?

I always find the Olympics to be an unusual experience. I’m hardly an athletics fanatic, yet I can’t help but get swept up in the spirit of the competition....

What is a vector database?

A vector database is a data storage system that organises information in the form of vectors, which are mathematical representations. These databases are...