卓越飞翔博客卓越飞翔博客

卓越飞翔 - 您值得收藏的技术分享站
技术文章35348本站已运行395

无法分解 Spark 数据框中的嵌套 JSON

无法分解 spark 数据框中的嵌套 json

问题内容

我是 spark 新手。我试图展平数据框,但未能通过“爆炸”做到这一点。

原始数据框架构如下:

id|approvaljson
1|[{"approvertype":"1st line manager","status":"approved"},{"approvertype":"2nd line manager","status":"approved"}]
2|[{"approvertype":"1st line manager","status":"approved"},{"approvertype":"2nd line manager","status":"rejected"}]

我需要将其转换为以下架构?

id|approvaltype|status
1|1st line manager|approved
1|2nd line manager|approved
2|1st line manager|approved
2|2nd line manager|rejected

我已经尝试过

df_exploded = df.withcolumn("approvaljson", explode("approvaljson"))

但是我得到了错误:


Cannot resolve "explode(ApprovalJSON)" due to data type mismatch:
parameter 1 requires ("ARRAY" or "MAP") type, however, "ApprovalJSON"
is of "STRING" type.;


正确答案


首先将类似 json 的字符串解析为结构数组,然后使用 inline 将数组分解为行和列

df1 = df.withcolumn("approvaljson", f.from_json("approvaljson", schema="array<struct>"))
df1 = df1.select("id", f.inline('approvaljson'))

结果

df1.show()

+---+----------------+--------+
| ID|    ApproverType|  Status|
+---+----------------+--------+
|  1|1st Line Manager|Approved|
|  1|2nd Line Manager|Approved|
|  2|1st Line Manager|Approved|
|  2|2nd Line Manager|Rejected|
+---+----------------+--------+
卓越飞翔博客
上一篇: ozzo 验证 v4 返回在结构中找不到字段 #0
下一篇: 返回列表
留言与评论(共有 0 条评论)
   
验证码:
隐藏边栏