Flink MySQL CDC connector 使用注意事项

最新推荐文章于 2025-06-07 09:04:26 发布

岚天逸剑

最新推荐文章于 2025-06-07 09:04:26 发布

阅读量267

点赞数

CC 4.0 BY-SA版权

分类专栏： flink hudi spark 文章标签： flink mysql hudi spark

本文链接：https://blog.youkuaiyun.com/Aquester/article/details/130628615

flink 同时被 3 个专栏收录

23 篇文章

订阅专栏

hudi

23 篇文章

订阅专栏

spark

19 篇文章

订阅专栏

文章指出了在使用HDFS时遇到的两个问题：一是表名不应包含大写字母，二是表名和库名不应包含点号，因为这可能导致查询失败和文件找不到的错误。问题在执行查询时暴露，如Java异常信息所示，系统尝试访问的大写路径与实际存在的小写路径不匹配。文章还提到了Flink写入和Spark读取时可能遇到的路径不一致问题。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

注意事项

表要有主键
库名和表名不能有点号

是个 BUG，估计后续会修复。

表名不能有大写

也是个 BUG，估计后续会修复。

如果表名含有大写的字母，查询时日志可看到如下信息：

java.util.concurrent.ExecutionException: java.io.FileNotFoundException: File does not exist: hdfs://hadoop/user/test/warehouse/test.db/ods_test
	at org.sparkproject.guava.util.concurrent.AbstractFuture.getDoneValue(AbstractFuture.java:552)
	at org.sparkproject.guava.util.concurrent.AbstractFuture.get(AbstractFuture.java:513)
	at org.sparkproject.guava.util.concurrent.AbstractFuture$TrustedFuture.get(AbstractFuture.java:90)
	at org.sparkproject.guava.util.concurrent.Uninterruptibles.getUninterruptibly(Uninterruptibles.java:199)
	at org.sparkproject.guava.cache.LocalCache$Segment.getAndRecordStats(LocalCache.java:2312)
	at org.sparkproject.guava.cache.LocalCache$Segment.loadSync(LocalCache.java:2278)
	at org.sparkproject.guava.cache.LocalCache$Segment.lockedGetOrLoad(LocalCache.java:2154)
	at org.sparkproject.guava.cache.LocalCache$Segment.get(LocalCache.java:2044)
	at org.sparkproject.guava.cache.LocalCache.get(LocalCache.java:3952)
	at org.sparkproject.guava.cache.LocalCache$LocalManualCache.get(LocalCache.java:4871)
	at org.apache.spark.sql.catalyst.catalog.SessionCatalog.getCachedPlan(SessionCatalog.scala:158)

而底层 HDFS 路径保持了大写：

2023-05-11 18:18:27,943 INFO  [47] [org.apache.hudi.util.StreamerUtil.initTableIfNotExists(StreamerUtil.java:335)]  - Table [hdfs://hadoop/user/test/warehouse/test.db/ODS_test/ODS_test] already exists, no need to initialize the table
2023-05-11 18:18:27,943 INFO  [47] [org.apache.hudi.util.StreamerUtil.initTableIfNotExists(StreamerUtil.java:337)]  - Table update under base path hdfs://hadoop/user/test/warehouse/test.db/ODS_test
2023-05-11 18:18:27,944 INFO  [47] [org.apache.hudi.common.table.HoodieTableMetaClient.updateTableAndGetMetaClient(HoodieTableMetaClient.java:507)]  - Update hoodie table with basePath hdfs://hadoop/user/test/warehouse/test.db/ODS_test
2023-05-11 18:18:27,949 INFO  [47] [org.apache.hudi.common.table.HoodieTableMetaClient.<init>(HoodieTableMetaClient.java:125)]  - Loading HoodieTableMetaClient from hdfs://hadoop/user/test/warehouse/test.db/ODS_test
2023-05-11 18:18:27,950 INFO  [47] [org.apache.hadoop.security.UserGroupInformation.initialize(UserGroupInformation.java:366)]  - Hadoop UGI authentication : TAUTH
2023-05-11 18:18:27,993 INFO  [47] [org.apache.hudi.common.table.HoodieTableConfig.<init>(HoodieTableConfig.java:295)]  - Loading table properties from hdfs://hadoop/user/test/warehouse/test.db/ODS_test/.hoodie/hoodie.properties
2023-05-11 18:18:28,212 INFO  [47] [org.apache.hudi.common.table.HoodieTableMetaClient.<init>(HoodieTableMetaClient.java:144)]  - Finished Loading Table of type MERGE_ON_READ(version=1, baseFileFormat=PARQUET) from hdfs://hadoop/user/test/warehouse/test.db/ODS_test
2023-05-11 18:18:29,730 INFO  [47] [org.apache.hudi.common.table.HoodieTableMetaClient.updateTableAndGetMetaClient(HoodieTableMetaClient.java:512)]  - Finished update Table of type MERGE_ON_READ from hdfs://hadoop/user/test/warehouse/test.db/ODS_test