簡體   English   中英

Bigquery中使用TIME格式時如何計算平均時間?

[英]How to calculate average time when it is used TIME format in Bigquery?

我正在嘗試獲取 AVG 時間,但 AVG function 不支持時間格式。 我嘗試使用 CAST function,就像在一些帖子中解釋的那樣,但它似乎無論如何都不起作用。 謝謝

WITH october_fall AS
   (SELECT
   start_station_name,
   end_station_name,
   start_station_id,
   end_station_id,
   EXTRACT (DATE FROM started_at) AS start_date,
   EXTRACT(DAYOFWEEK FROM started_at) AS start_week_date,
   EXTRACT (TIME FROM started_at) AS start_time,    
   EXTRACT (DATE FROM ended_at) AS end_date,
   EXTRACT(DAYOFWEEK FROM ended_at) AS end_week_date,    
   EXTRACT (TIME FROM ended_at) AS end_time,
   DATETIME_DIFF (ended_at,started_at, MINUTE) AS total_lenght,
   member_casual
FROM 
   `ciclystic.cyclistic_seasonal_analysis.fall_202010` AS fall_analysis
ORDER BY 
   started_at DESC)
SELECT
   COUNT (start_week_date) AS avg_start_1,
   AVG (start_time) AS avg_start_time_1, ## here is where the problem start
   member_casual
FROM 
   october_fall
WHERE 
   start_week_date = 1
GROUP BY
   member_casual

由於 BigQuery 無法計算 TIME 類型的 AVG,如果您嘗試這樣做,您會看到錯誤消息。

相反,您可以通過 INT64 計算 AVG。
time_ts是時間戳格式。
我嘗試使用time_diff計算從時間到“00:00:00”的差異,然后我可以獲得 FLOAT64 格式的秒數並將其轉換為 INT64 格式。
我創建了一個函數secondToTime 計算小時/分鍾/秒並解析回時間格式非常簡單。

對於日期格式,我認為你可以用同樣的方式來做。

create temp function secondToTime (seconds INT64)
    returns time 
    as (
        PARSE_TIME (
            "%H:%M:%S",
            concat(
                cast(seconds / 3600 as int),
                ":",
                cast(mod(seconds, 3600) / 60 as int),
                ":",
                mod(seconds, 60)
            )
        )
    );


with october_fall as (
    select
        extract (date from time_ts) as start_date,
        extract (time from time_ts) as start_time
    from `bigquery-public-data.hacker_news.comments`
    limit 10
) SELECT 
    avg(time_diff(start_time, time '00:00:00', second)),
    secondToTime(
        cast(avg(time_diff(start_time, time '00:00:00', second)) as INT64) 
    ),
    secondToTime(0),
    secondToTime(60),
    secondToTime(3601),
    secondToTime(7265)
FROM october_fall

試試下面

SELECT
   COUNT (start_week_date) AS avg_start_1,
   TIME(
     EXTRACT(hour   FROM AVG(start_time - '0:0:0')), 
     EXTRACT(minute FROM AVG(start_time - '0:0:0')), 
     EXTRACT(second FROM AVG(start_time - '0:0:0'))
   ) as avg_start_time_1
   member_casual
FROM 
   october_fall
WHERE 
   start_week_date = 1
GROUP BY
   member_casual     

另一種選擇是

SELECT
   COUNT (start_week_date) AS avg_start_1,
   PARSE_TIME('0-0 0 %H:%M:%E*S', '' || AVG(start_time - '0:0:0')) as avg_start_time_1
   member_casual
FROM 
   october_fall
WHERE 
   start_week_date = 1
GROUP BY
   member_casual     

我知道幾個月過去了,但也許其他人會面臨同樣的問題。 至於出現問題的部分,這樣的事情對我有用,並給出了平均ride_length:

FORMAT_TIMESTAMP
  ('%T', 
  TIMESTAMP_SECONDS(CAST(AVG(TIME_DIFF(ride_length, '00:00:00', SECOND)) AS 
  INT64)))
   AS avg_ride_length

暫無
暫無

聲明:本站的技術帖子網頁,遵循CC BY-SA 4.0協議,如果您需要轉載,請注明本站網址或者原文地址。任何問題請咨詢:yoyou2525@163.com.

 
粵ICP備18138465號  © 2020-2024 STACKOOM.COM