将Apache Spark Scala代码转换为Python

Question

Can anyone convert this very simple scala code to python? 任何人都可以将这个非常简单的scala代码转换为python吗？

val words = Array("one", "two", "two", "three", "three", "three")
val wordPairsRDD = sc.parallelize(words).map(word => (word, 1))

val wordCountsWithGroup = wordPairsRDD
    .groupByKey()
    .map(t => (t._1, t._2.sum))
    .collect()

Answer 1

try this: 尝试这个：

words = ["one", "two", "two", "three", "three", "three"]
wordPairsRDD = sc.parallelize(words).map(lambda word : (word, 1))

wordCountsWithGroup = wordPairsRDD
    .groupByKey()
    .map(lambda t: (t[0], sum(t[1])))
    .collect()

Answer 2

Two translate in python : 两个在python中翻译：

from operator import add
wordsList = ["one", "two", "two", "three", "three", "three"]
words = sc.parallelize(wordsList ).map(lambda l :(l,1)).reduceByKey(add).collect()
print words
words = sc.parallelize(wordsList ).map(lambda l : (l,1)).groupByKey().map(lambda t: (t[0], sum(t[1]))).collect()
print words

Answer 3

Assuming you already have a Spark context defined and ready to go: 假设您已经定义了Spark上下文并准备好了：

 from operator import add
 words = ["one", "two", "two", "three", "three", "three"]
 wordsPairRDD = sc.parallelize(words).map(lambda word: (word, 1))
      .reduceByKey(add)
      .collect()

Checkout the github examples repo: Python Examples 查看github示例repo： Python示例

将Apache Spark Scala代码转换为Python

问题描述

3 个解决方案

解决方案1
5 已采纳 2015-06-12 20:49:40

解决方案2
2 2015-06-12 20:53:52

解决方案3
2 2015-06-12 20:58:05

将Apache Spark Scala代码转换为Python

问题描述

3 个解决方案

解决方案1 5 已采纳 2015-06-12 20:49:40

解决方案2 2 2015-06-12 20:53:52

解决方案3 2 2015-06-12 20:58:05

解决方案1
5 已采纳 2015-06-12 20:49:40

解决方案2
2 2015-06-12 20:53:52

解决方案3
2 2015-06-12 20:58:05