combine utf-16 surrogate pairs in json string parser
What changed, and why it matters
This commit fixes how Monero's JSON parser handles special Unicode escape sequences called UTF-16 surrogate pairs. Previously, the parser treated each half of a surrogate pair as a separate character, which could produce invalid UTF-8 output or fail to decode characters outside the basic multilingual plane correctly. The patch now combines the two halves into the correct single Unicode character and rejects malformed or lone surrogates. This is a correctness fix in JSON string parsing; it does not appear to be a critical security vulnerability on its own, but malformed Unicode handling can sometimes lead to downstream issues in systems that process the parsed strings.
Apply the patch to ensure JSON strings containing non-BMP Unicode characters are parsed correctly and invalid surrogate pairs are rejected. Review any downstream components that may have received malformed UTF-8 from earlier versions, though no specific exploit is indicated by the commit alone.
Security signals we found
JSON parser Unicode handling correction
UTF-16 surrogate pair validation added
Rejection of lone/malformed surrogates
Potential for invalid UTF-8 output before patch
Evidence from the diff
The change is in contrib/epee/src/parserse_base_utils.cpp, which implements JSON string unescaping. Before the patch, \uXXXX escapes were decoded individually and converted to UTF-8 without recognizing UTF-16 surrogate pairs. The new code detects high surrogates (0xD800-0xDBFF), requires a following \uXXXX low surrogate (0xDC00-0xDFFF), combines them into a code point in the U+10000 to U+10FFFF range using the standard formula, and rejects lone or mismatched surrogates. It also extends the UTF-8 encoder to emit the correct 4-byte sequence for code points above U+FFFF. Unit tests verify valid surrogate pair decoding and rejection of invalid cases.
Changed components
contrib/epee/src/parserse_base_utils.cppepee JSON deserialization pathInspect captured patch +68 / −11
diff --git a/contrib/epee/src/parserse_base_utils.cpp b/contrib/epee/src/parserse_base_utils.cpp
index 923dde4..5dbd03f 100644
--- a/contrib/epee/src/parserse_base_utils.cpp
+++ b/contrib/epee/src/parserse_base_utils.cpp
@@ -129,18 +129,31 @@ namespace misc_utils
case '/': //Slash character
val.push_back('/');break;
case 'u': //Unicode code point
- if (buf_end - it < 5)
{
- ASSERT_MES_AND_THROW("Invalid Unicode escape sequence");
- }
- else
- {
- uint32_t dst = 0;
- for (int i = 0; i < 4; ++i)
+ auto read_hex4 = [&]() -> uint32_t {
+ CHECK_AND_ASSERT_THROW_MES(buf_end - it >= 5, "Invalid Unicode escape sequence");
+ uint32_t v = 0;
+ for (int i = 0; i < 4; ++i)
+ {
+ const unsigned char tmp = isx[(unsigned char)*++it];
+ CHECK_AND_ASSERT_THROW_MES(tmp != 0xff, "Bad Unicode encoding");
+ v = v << 4 | tmp;
+ }
+ return v;
+ };
+ uint32_t dst = read_hex4();
+ // combine a UTF-16 surrogate pair into a single code point; reject lone surrogates
+ if (dst >= 0xd800 && dst <= 0xdbff)
{
- const unsigned char tmp = isx[(unsigned char)*++it];
- CHECK_AND_ASSERT_THROW_MES(tmp != 0xff, "Bad Unicode encoding");
- dst = dst << 4 | tmp;
+ CHECK_AND_ASSERT_THROW_MES(buf_end - it >= 3 && *(it + 1) == '\\' && *(it + 2) == 'u', "Invalid UTF-16 surrogate pair");
+ it += 2;
+ const uint32_t low = read_hex4();
+ CHECK_AND_ASSERT_THROW_MES(low >= 0xdc00 && low <= 0xdfff, "Invalid UTF-16 surrogate pair");
+ dst = 0x10000 + ((dst - 0xd800) << 10) + (low - 0xdc00);
+ }
+ else
+ {
+ CHECK_AND_ASSERT_THROW_MES(dst < 0xdc00 || dst > 0xdfff, "Invalid UTF-16 surrogate pair");
}
// encode as UTF-8
if (dst <= 0x7f)
@@ -160,7 +173,10 @@ namespace misc_utils
}
else
{
- ASSERT_MES_AND_THROW("Unicode code point is out or range");
+ val.push_back(0xf0 | (dst >> 18));
+ val.push_back(0x80 | ((dst >> 12) & 0x3f));
+ val.push_back(0x80 | ((dst >> 6) & 0x3f));
+ val.push_back(0x80 | (dst & 0x3f));
}
}
break;
diff --git a/tests/unit_tests/epee_serialization.cpp b/tests/unit_tests/epee_serialization.cpp
index 2cafc0e..0502ab6 100644
--- a/tests/unit_tests/epee_serialization.cpp
+++ b/tests/unit_tests/epee_serialization.cpp
@@ -202,3 +202,44 @@ TEST(epee_binary, any_empty_seq)
EXPECT_TRUE(epee::serialization::load_t_from_binary(i, epee::span<const std::uint8_t>(data_empty_object)));
EXPECT_EQ(0, i.x.size());
}
+
+namespace
+{
+struct ObjOfString
+{
+ std::string x;
+
+ BEGIN_KV_SERIALIZE_MAP()
+ KV_SERIALIZE(x)
+ END_KV_SERIALIZE_MAP()
+};
+}
+
+TEST(epee_json, unicode_escape_surrogate_pair)
+{
+ // 😀 is the UTF-16 surrogate pair for U+1F600, which must decode to
+ // the 4-byte UTF-8 sequence F0 9F 98 80 (not two 3-byte encoded surrogates).
+ ObjOfString o{};
+ EXPECT_TRUE(epee::serialization::load_t_from_json(o, "{\"x\":\"\\uD83D\\uDE00\"}"));
+ EXPECT_EQ(std::string("\xF0\x9F\x98\x80"), o.x);
+
+ // Highest valid code point U+10FFFF.
+ ObjOfString o2{};
+ EXPECT_TRUE(epee::serialization::load_t_from_json(o2, "{\"x\":\"\\uDBFF\\uDFFF\"}"));
+ EXPECT_EQ(std::string("\xF4\x8F\xBF\xBF"), o2.x);
+
+ // A basic multilingual plane escape is unaffected.
+ ObjOfString o3{};
+ EXPECT_TRUE(epee::serialization::load_t_from_json(o3, "{\"x\":\"\\u20AC\"}"));
+ EXPECT_EQ(std::string("\xE2\x82\xAC"), o3.x);
+}
+
+TEST(epee_json, unicode_escape_bad_surrogate)
+{
+ // A lone high surrogate, a lone low surrogate, and a high surrogate not
+ // followed by a low surrogate are all invalid and must fail to parse.
+ ObjOfString o{};
+ EXPECT_FALSE(epee::serialization::load_t_from_json(o, "{\"x\":\"\\uD83D\"}"));
+ EXPECT_FALSE(epee::serialization::load_t_from_json(o, "{\"x\":\"\\uDE00\"}"));
+ EXPECT_FALSE(epee::serialization::load_t_from_json(o, "{\"x\":\"\\uD83D\\u0041\"}"));
+}
Why this scored 51/100
Community notes
Notes can correct, qualify, or add evidence to the AI analysis. Every note shown here has been validated by a human moderator.
The AI analysis stands alone for now. Submit a note if you can add evidence or important context.