简体   繁体   English

如何从NVARCHAR(MAX)属性解析编码为UTF-8的XML?

[英]How to parse XML encoded as UTF-8 from a NVARCHAR(MAX) attribute?

I'm facing a problem to parse an XML string stored in a field of type NVARCHAR(MAX) (I cannot change the type of this field). 我在解析存储在NVARCHAR(MAX)类型字段中的XML字符串时遇到问题(我无法更改此字段的类型)。

Here is my table (WorkingHours) : 这是我的桌子(WorkingHours):

CREATE TABLE WorkingHours(
    [ID] [int] NOT NULL PRIMARY KEY,
    [CONTENT] [nvarchar](MAX) NOT NULL,
    -- ...
);

Here is a sample of the [CONTENT] attribute : 以下是[CONTENT]属性的示例:

<?xml version="1.0" encoding="UTF-8"?>
    <calendar>
        <day number="1" worked_day="no">
            <interval number="1" begin_hour="08:30" end_hour="12:00"/>
            <interval number="2" begin_hour="13:30" end_hour="17:00"/>
            <interval number="3" begin_hour="" end_hour=""/></day>
        <day number="2" worked_day="no">
            <interval number="1" begin_hour="08:30" end_hour="12:00"/>
            <interval number="2" begin_hour="13:30" end_hour="17:00"/>
            <interval number="3" begin_hour="" end_hour=""/>
        </day>
        <day number="3" worked_day="no">
            <interval number="1" begin_hour="08:30" end_hour="12:00"/>
            <interval number="2" begin_hour="13:30" end_hour="17:00"/>
            <interval number="3" begin_hour="" end_hour=""/>
        </day>
        <day number="4" worked_day="no">
            <interval number="1" begin_hour="08:30" end_hour="12:00"/>
            <interval number="2" begin_hour="13:30" end_hour="17:00"/>
            <interval number="3" begin_hour="" end_hour=""/>
        </day>
        <day number="5" worked_day="no">
            <interval number="1" begin_hour="08:30" end_hour="12:00"/>
            <interval number="2" begin_hour="13:30" end_hour="17:00"/>
            <interval number="3" begin_hour="" end_hour=""/>
        </day>
        <day number="6" worked_day="no">
            <interval number="1" begin_hour="" end_hour=""/>
            <interval number="2" begin_hour="" end_hour=""/>
            <interval number="3" begin_hour="" end_hour=""/>
        </day>
        <day number="7" worked_day="no">
            <interval number="1" begin_hour="" end_hour=""/>
            <interval number="2" begin_hour="" end_hour=""/>
            <interval number="3" begin_hour="" end_hour=""/>
        </day>
    </calendar>

As you can see, the data encoding is UTF-8 . 如您所见,数据编码为UTF-8

Now, I would like to parse this data in order to create some calculations : 现在,我想解析这些数据以创建一些计算:

DECLARE @RawContent [nvarchar](MAX) = (
    SELECT wh.[CONTENT]
    FROM [WorkingHours] wh 
    WHERE wh.[ID] = 100);

DECLARE @XMLContent [Xml] = @RawContent; // KO
-- DECLARE @XMLContent [Xml] = CAST(@RawContent AS XML);  // KO
-- DECLARE @XMLContent [Xml] = CONVERT(XML, @RawContent); // KO

-- Just a test to query XML data.
SELECT 
    C.WD.value('@number', 'int') AS DayId         
FROM @XMLContent.nodes('/calendar/day') AS C(WD);   

I don't know how to cast the result (a nvarchar(max) field containing UTF-8 XML string) to a XML value. 我不知道如何将结果(包含UTF-8 XML字符串的nvarchar(max)字段)转换为XML值。 SQL Server returns the following error : SQL Server返回以下错误:

"Unable to switch encoding"

It refers to the CAST line (when I define the @XMLContent variable). 它指的是CAST行(当我定义@XMLContent变量时)。

Any idea to solve that ? 有什么想法解决这个问题?

Strip out the processing directive -- it's meaningless and incorrect because the data is already encoded in UTF-16 (since it's stored as NVARCHAR ). 删除处理指令 - 它没有意义且不正确,因为数据已经以UTF-16编码(因为它存储为NVARCHAR )。 If you cannot change the data already present, you'll have to rely on (slightly brittle) string replacement: 如果您无法更改已存在的数据,则必须依赖(略微脆弱)字符串替换:

CAST(REPLACE(wh.[CONTENT], '<?xml version="1.0" encoding="UTF-8"?>', '') AS XML)

Note that explicitly indicating the encoding is UTF-16 instead will also work -- though it adds nothing. 请注意,显式指示编码是UTF-16也可以工作 - 虽然它什么都不添加。

The other option is to convert to a VARCHAR datatype first - which is non-Unicode - and then to XML : 另一种选择是首先转换为VARCHAR数据类型 - 非Unicode - 然后转换为XML

DECLARE @RawContent [nvarchar](MAX) = (
    SELECT wh.[CONTENT]
    FROM [WorkingHours] wh 
    WHERE wh.[ID] = 100);

DECLARE @XMLContent XML = CAST(CAST(@RawContent AS VARCHAR(MAX)) AS XML)

-- Just a test to query XML data.
SELECT 
    C.WD.value('@number', 'int') AS DayId         
FROM @XMLContent.nodes('/calendar/day') AS C(WD);   

声明:本站的技术帖子网页,遵循CC BY-SA 4.0协议,如果您需要转载,请注明本站网址或者原文地址。任何问题请咨询:yoyou2525@163.com.

 
粤ICP备18138465号  © 2020-2024 STACKOOM.COM