简体   繁体   English

连接 Athena 中的变量 SQL 查询来自 Python Lambda ZC1C425268E68385D14ZA5074CF

[英]Concatenate variable in Athena SQL query from Python Lambda function

I have a Python Lambda function that creates a SQL table in Athena.我有一个 Python Lambda function 可以在 AZA2A21 中创建一个 Z9778840A0100CB30C98287676741 表。 How do I properly concatenate variables in my query?如何正确连接查询中的变量? When I set the LOCATION value, I receive the error response below.当我设置 LOCATION 值时,我收到下面的错误响应。 The function runs successfully if I hard code the LOCATION value.如果我对 LOCATION 值进行硬编码,则 function 会成功运行。

LOCATION “”” + s3_bucket_test + “””

Error response:错误响应:

Response
{
  "errorMessage": "An error occurred (InvalidRequestException) when calling the StartQueryExecution operation: line 1:8: mismatched input 'EXTERNAL'. Expecting: 'OR', 'SCHEMA', 'TABLE', 'VIEW'",
  "errorType": "InvalidRequestException",
  "stackTrace": [
    "  File \"/var/task/lambda_function.py\", line 34, in lambda_handler\n    queryStart = client.start_query_execution(\n",
    "  File \"/var/runtime/botocore/client.py\", line 386, in _api_call\n    return self._make_api_call(operation_name, kwargs)\n",
    "  File \"/var/runtime/botocore/client.py\", line 705, in _make_api_call\n    raise error_class(parsed_response, operation_name)\n"
  ]
}

Lambda function: Lambda function:

import boto3
import json
import time

database = ‘daily_reports’
s3_bucket = 's3://test/’
s3_bucket_results = 's3://test/results’

query = ("""
          CREATE EXTERNAL TABLE IF NOT EXISTS `reports` (
              `timestamp` bigint,
              `user_id` string,
              `name` string
            )
            ROW FORMAT SERDE 'org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe' 
            WITH SERDEPROPERTIES (
              'serialization.format' = '1'
            ) LOCATION “”” + s3_bucket + “””
            TBLPROPERTIES ('has_encrypted_data'='false');
        """)        


def lambda_handler(event, context):
    client = boto3.client('athena')
    queryStart = client.start_query_execution(
        QueryString = query,
        QueryExecutionContext = {
            'Database': database
        },
        ResultConfiguration = {
            'OutputLocation': s3_bucket_results
        }
    )

Thank you.谢谢你。

Have you tried to use Python's format method?你试过用 Python 的 format 方法吗? Something like this像这样的东西

query = ("""
          CREATE EXTERNAL TABLE IF NOT EXISTS `reports` (
              `timestamp` bigint,
              `user_id` string,
              `name` string
            )
            ROW FORMAT SERDE 'org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe' 
            WITH SERDEPROPERTIES (
              'serialization.format' = '1'
            ) LOCATION '{}' 
            TBLPROPERTIES ('has_encrypted_data'='false');
        """).format(s3_bucket)    

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM