SPARK SQL - update MySql table using DataFrames and JDBC

最新推荐文章于 2025-01-06 17:00:00 发布

转载最新推荐文章于 2025-01-06 17:00:00 发布 · 3k 阅读

Spark 专栏收录该内容

55 篇文章

订阅专栏

本文探讨了使用Spark SQL DataFrames及JDBC连接在MySQL中插入及更新数据的方法。介绍了如何通过不同SaveMode进行操作，并提出了实现类似MySQL 'ON DUPLICATE KEY UPDATE' 功能的几种替代方案。

I'm trying to insert and update some data on MySql using Spark SQL DataFrames and JDBC connection.

I've succeeded to insert new data using the SaveMode.Append. Is there a way to update the data already existing in MySql Table from Spark SQL?

My code to insert is:

myDataFrame.write.mode(SaveMode.Append).jdbc(JDBCurl,mySqlTable,connectionProperties)

If I change to SaveMode.Overwrite it deletes the full table and creates a new one, I'm looking for something like the "ON DUPLICATE KEY UPDATE" available in MySql

It is not possible. As for now (Spark 1.6.0 / 2.2.0 SNAPSHOT) Spark DataFrameWriter supports only four writing modes:

SaveMode.Overwrite: overwrite the existing data.
SaveMode.Append: append the data.
SaveMode.Ignore: ignore the operation (i.e. no-op).
SaveMode.ErrorIfExists: default option, throw an exception at runtime.

You can insert manually for example using mapPartitions (since you want an UPSERT operation should be idempotent and as such easy to implement), write to temporary table and execute upsert manually, or use triggers.

In general achieving upsert behavior for batch operations and keeping decent performance is far from trivial. You have to remember that in general case there will be multiple concurrent transactions in place (one per each partition) so you have to ensure that there will no write conflicts (typically by using application specific partitioning) or provide appropriate recovery procedures. In practice it may be better to perform and batch writes to a temporary table and resolve upsert part directly in the database.