需要一个从yaml文件内容提取并作为csv文件输出的脚本

Question

我是python的新手，但很感谢您的指导，帮助我创建了一个简单的脚本，该脚本读取一堆.yaml文件（同一目录中约有300个文件），并从中提取了某个部分（仅限选修科目） .yaml文件并将其转换为csv。

.yaml文件中的内容的一个示例

code: 9313
degrees:
- name: Design
  coreCourses:
  - ABCD1
  - ABCD2
  - ABCD3
  electiveGroups: #this is the section i need to extract
    - label: Electives
      options:
        - Studio1
        - Studio2
        - Studio3
    - label: OtherElectives
      options:
        - Class1
        - Development2
        - lateclass1
   specialisations:
    - label: Honours

我想如何查看csv中的输出：

.yaml file name | Electives   | Studio1
.yaml file name | Electives   | Studio2
.yaml file name | Electives   | Studio3
.yaml file name | OtherElectives   | class1
.yaml file name | OtherElectives   | Development2
.yaml file name | OtherElectives   | lateclass1

我假设这将是一个相对简单的脚本，但是我正在寻找一些帮助来编写此脚本。 我对此很陌生，所以请耐心等待。 我已经写了一些vba宏，所以我希望我可以相对较快地掌握。

最好的办法是提供完整的解决方案，并提供有关代码工作方式的一些指导。

感谢您提前提供的所有帮助。 我希望我的问题很清楚

这是我的第一次尝试（尽管花了很长时间）：

import yaml
with open ('program_4803','r') as f:
    doc = yaml.load(f)
    txt=doc["electiveGroups"]["options"]
    file = open(“test.txt”,”w”) 
        file.write(“txt”) 
        file.close()

您可能会说，目前这还很不完整-但我正在尽我最大的努力！

Answer 1

要解析yaml文件，请使用python yaml库

此处的示例：在Python中解析YAML文件并访问数据？

要写入文件，您不需要csv库

file = open(“testfile.txt”,”w”) 
file.write(“Hello World”) 
file.close()

上面的代码将写入文件，您可以仅迭代yaml解析的结果，然后将输出相应地写入文件。

Answer 2

这可能会有所帮助：

import yaml
import csv

yaml_file_names = ['data.yaml', 'data2.yaml']


rows_to_write = []

for idx, each_yaml_file in enumerate(yaml_file_names):
    print("Processing file ", idx+1, "of", len(yaml_file_names), "file name:", each_yaml_file)
    with open(each_yaml_file) as f:
        data = yaml.load(f)

        for each_dict in data['degrees']:
            for each_nested_dict in each_dict['electiveGroups']:
                for each_option in each_nested_dict['options']:
                    # write to csv yaml_file_name, each_nested_dict['label'], each_option
                    rows_to_write.append([each_yaml_file, each_nested_dict['label'], each_option])



with open('output_csv_file.csv', 'w') as out:
    csv_writer = csv.writer(out, delimiter='|')
    csv_writer.writerows(rows_to_write)
    print("Output file output_csv_file.csv created")

使用两个模拟输入yaml的data.yaml和data2.yaml测试了此代码，其内容如下：

data.yaml ：

code: 9313
degrees:
- name: Design
  coreCourses:
  - ABCD1
  - ABCD2
  - ABCD3
  electiveGroups: #this is the section i need to extract
    - label: Electives
      options:
        - Studio1
        - Studio2
        - Studio3
    - label: OtherElectives
      options:
        - Class1
        - Development2
        - lateclass1
  specialisations:
  - label: Honours

和data2.yaml ：

code: 9313
degrees:
- name: Design
  coreCourses:
  - ABCD1
  - ABCD2
  - ABCD3
  electiveGroups: #this is the section i need to extract
    - label: Electives
      options:
        - Studio1
    - label: E2
      options:
        - Class1
  specialisations:
  - label: Honours

并且生成的输出csv文件是这样的：

data.yaml|Electives|Studio1
data.yaml|Electives|Studio2
data.yaml|Electives|Studio3
data.yaml|OtherElectives|Class1
data.yaml|OtherElectives|Development2
data.yaml|OtherElectives|lateclass1
data2.yaml|Electives|Studio1
data2.yaml|E2|Class1

顺便说一句，您输入的Yaml输入以及您的问题，最后两行未正确缩进

正如您所说的那样，您需要解析目录中的300个yaml文件，那么您可以使用python的glob模块，如下所示：

import yaml
import csv
import glob


yaml_file_names = glob.glob('./*.yaml')
# yaml_file_names = ['data.yaml', 'data2.yaml']

rows_to_write = []

for idx, each_yaml_file in enumerate(yaml_file_names):
    print("Processing file ", idx+1, "of", len(yaml_file_names), "file name:", each_yaml_file)
    with open(each_yaml_file) as f:
        data = yaml.load(f)

        for each_dict in data['degrees']:
            for each_nested_dict in each_dict['electiveGroups']:
                for each_option in each_nested_dict['options']:
                    # write to csv yaml_file_name, each_nested_dict['label'], each_option
                    rows_to_write.append([each_yaml_file, each_nested_dict['label'], each_option])



with open('output_csv_file.csv', 'w') as out:
    csv_writer = csv.writer(out, delimiter='|', quotechar=' ')
    csv_writer.writerows(rows_to_write)
    print("Output file output_csv_file.csv created")

编辑：如您在注释中所要求的，以跳过那些没有electiveGroup部分的yaml文件，这是更新的程序：

import yaml
import csv
import glob


yaml_file_names = glob.glob('./*.yaml')
# yaml_file_names = ['data.yaml', 'data2.yaml']

rows_to_write = []

for idx, each_yaml_file in enumerate(yaml_file_names):
    print("Processing file ", idx+1, "of", len(yaml_file_names), "file name:", each_yaml_file)
    with open(each_yaml_file) as f:
        data = yaml.load(f)

        for each_dict in data['degrees']:
            try:
                for each_nested_dict in each_dict['electiveGroups']:
                    for each_option in each_nested_dict['options']:
                        # write to csv yaml_file_name, each_nested_dict['label'], each_option
                        rows_to_write.append([each_yaml_file, each_nested_dict['label'], each_option])
            except KeyError:
                print("No electiveGroups or options key found in", each_yaml_file)


with open('output_csv_file.csv', 'w') as out:
    csv_writer = csv.writer(out, delimiter='|', quotechar=' ')
    csv_writer.writerows(rows_to_write)
    print("Output file output_csv_file.csv created")

需要一个从yaml文件内容提取并作为csv文件输出的脚本

问题描述

2 个解决方案

解决方案1
0 2017-10-11 04:34:19

解决方案2
0 已采纳 2017-10-11 05:19:44

需要一个从yaml文件内容提取并作为csv文件输出的脚本

问题描述

2 个解决方案

解决方案1 0 2017-10-11 04:34:19

解决方案2 0 已采纳 2017-10-11 05:19:44

解决方案1
0 2017-10-11 04:34:19

解决方案2
0 已采纳 2017-10-11 05:19:44