简体   繁体   English

无法在本地连接S3和Spark

[英]Unable to connect with S3 and Spark Locally

Below is my code : I am trying to access s3 files from spark locally. 下面是我的代码:我正在尝试从本地访问Spark的s3文件。 But getting error : Exception in thread "main" org.apache.hadoop.security.AccessControlException: Permission denied: s3n://bucketname/folder I am also using jars :hadoop-aws-2.7.3.jar,aws-java-sdk-1.7.4.jar,hadoop-auth-2.7.1.jar while submitting spark job from cmd. 但是出现错误:线程“ main” org.apache.hadoop.security.AccessControlException中的异常:权限被拒绝:s3n:// bucketname / folder我也在使用jars:hadoop-aws-2.7.3.jar,aws-java- sdk-1.7.4.jar,hadoop-auth-2.7.1.jar,同时从cmd提交spark作业。

package org.test.snow
import org.apache.spark._
import org.apache.spark.SparkContext._
import org.apache.log4j._
import org.apache.spark.storage.StorageLevel
import org.apache.spark.sql.SparkSession
import org.apache.spark.util.Utils
import org.apache.spark.sql._
import org.apache.hadoop.fs.FileSystem
import org.apache.hadoop.fs.Path

object SnowS3 {
def main(args: Array[String]) {
val conf = new SparkConf().setAppName("IDV4")
val sc = new SparkContext(conf)
val spark = new org.apache.spark.sql.SQLContext(sc)
import spark.implicits._
sc.hadoopConfiguration.set("fs.s3a.impl","org.apache.hadoop.fs.s3native.NativeS3FileSystem")
sc.hadoopConfiguration.set("fs.s3a.awsAccessKeyId", "A*******************A")
sc.hadoopConfiguration.set("fs.s3a.awsSecretAccessKey","A********************A")
val cus_1=spark.read.format("com.databricks.spark.csv")
.option("header","true")
.option("inferSchema","true")
.load("s3a://tb-us-east/working/customer.csv")
cus_1.show()
    }
}

Any help would be appreciated. 任何帮助,将不胜感激。 FYI: I am using spark 2.1 仅供参考:我正在使用spark 2.1

You shouldn't set that fs.s3a.impl option; 您不应该设置fs.s3a.impl选项。 that's a superstition which seems to persist in spark examples. 这是一种迷信,似乎在火花示例中仍然存在。

Instead uses the S3A connector just by using the s3a:// prefix with 而是仅通过使用带有s3a://前缀的S3A连接器

  • consistent versions of hadoop-* jar versions. hadoop- * jar版本的一致版本。 Yes, hadoop-aws-2.7.3 needs hadoop-common-2.7.3 是的,hadoop-aws-2.7.3需要hadoop-common-2.7.3
  • setting the s3a specific authentication options, fs.s3a.access.key and `fs.s3a.secret.key' 设置s3a特定的身份验证选项fs.s3a.access.key和`fs.s3a.secret.key'

If that doesn't work, look at the s3a troubleshooting docs 如果这不起作用,请查看s3a故障排除文档

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM