функция since 6.9.0

wp_scrub_utf8()

Проверено на WordPress 6.9, обновлено Источник: WordPress Developer Resources.

Сигнатура

wp_scrub_utf8( string $text ): string

Описание

Понять, что делать при проблемах с кодировкой текста, бывает непросто.
Эта функция заменяет некорректные участки байтов, чтобы нейтрализовать возможные повреждения и не дать им вызвать дальнейшие проблемы в последующей обработке.
Однако заменять эти байты не всегда уместно. В некоторых случаях лучше оставить некорректные байты в строке, чтобы последующий код мог обработать их особым образом. Слишком ранняя замена байтов, как и слишком раннее экранирование для HTML, может привести к другим видам повреждения и потере данных.
В случае сомнений используйте эту функцию для замены участков некорректных байтов.
Замена выполняется по алгоритму «максимальной подчасти» для безопасных и совместимых строк. Это может приводить к нескольким символам замены подряд.
Пример:

// Valid strings come through unchanged.
'test' === wp_scrub_utf8( 'test' );

// Invalid sequences of bytes are replaced.
$invalid = "the byte xC0 is never allowed in a UTF-8 string.";
"the byte \u{FFFD} is never allowed in a UTF-8 string." === wp_scrub_utf8( $invalid, true );
'the byte � is never allowed in a UTF-8 string.' === wp_scrub_utf8( $invalid, true );

// Maximal subparts are replaced individually.
'.�.' === wp_scrub_utf8( ".\xC0." ); // C0 is never valid.
'.�.' === wp_scrub_utf8( ".\xE2\x8C." ); // Missing A3 at end.
'.��.' === wp_scrub_utf8( ".\xE2\x8C\xE2\x8C." ); // Maximal subparts replaced separately.
'.��.' === wp_scrub_utf8( ".\xC1\xBF." ); // Overlong sequence.
'.���.' === wp_scrub_utf8( ".\xED\xA0\x80." ); // Surrogate half.Внимание! Символ замены Unicode сам является символом Unicode (U+FFFD).
После того как участок некорректных байтов заменён им, невозможно определить, был ли символ замены изначально задуман или он появился в результате очистки байтов. Идеально оставлять замену только для отображения, но некоторые контексты (например, генерация XML или передача данных в большую языковую модель) требуют корректных входных строк.
См. такжеhttps://www.unicode.org/versions/Unicode16.0.0/core-spec/chapter-5/#G40630

Оригинал (английский)

Knowing what to do in the presence of text encoding issues can be complicated.
This function replaces invalid spans of bytes to neutralize any corruption that may be there and prevent it from causing further problems downstream.

However, it’s not always ideal to replace those bytes. In some settings it may be best to leave the invalid bytes in the string so that downstream code can handle them in a specific way. Replacing the bytes too early, like escaping for HTML too early, can introduce other forms of corruption and data loss.

When in doubt, use this function to replace spans of invalid bytes.

Replacement follows the “maximal subpart” algorithm for secure and interoperable strings. This can lead to sequences of multiple replacement characters in a row.

Example:

// Valid strings come through unchanged.
'test' === wp_scrub_utf8( 'test' );

// Invalid sequences of bytes are replaced.
$invalid = "the byte xC0 is never allowed in a UTF-8 string.";
"the byte \u{FFFD} is never allowed in a UTF-8 string." === wp_scrub_utf8( $invalid, true );
'the byte � is never allowed in a UTF-8 string.' === wp_scrub_utf8( $invalid, true );

// Maximal subparts are replaced individually.
'.�.' === wp_scrub_utf8( ".\xC0." );              // C0 is never valid.
'.�.' === wp_scrub_utf8( ".\xE2\x8C." );          // Missing A3 at end.
'.��.' === wp_scrub_utf8( ".\xE2\x8C\xE2\x8C." ); // Maximal subparts replaced separately.
'.��.' === wp_scrub_utf8( ".\xC1\xBF." );         // Overlong sequence.
'.���.' === wp_scrub_utf8( ".\xED\xA0\x80." );    // Surrogate half.

Note! The Unicode Replacement Character is itself a Unicode character (U+FFFD).
Once a span of invalid bytes has been replaced by one, it will not be possible to know whether the replacement character was originally intended to be there or if it is the result of scrubbing bytes. It is ideal to leave replacement for display only, but some contexts (e.g. generating XML or passing data into a large language model) require valid input strings.

See also

Параметры

$text string обязательный
Строка, которая предположительно является UTF-8, но может содержать некорректные последовательности байтов.

Возвращаемое значение

string

Исходный код

wp-includes/utf8.php:109

function wp_scrub_utf8( $text ) {
	/*
	 * While it looks like setting the substitute character could fail,
	 * the internal PHP code will never fail when provided a valid
	 * code point as a number. In this case, theres no need to check
	 * its return value to see if it succeeded.
	 */
	$prev_replacement_character = mb_substitute_character();
	mb_substitute_character( xFFFD );
	$scrubbed = mb_scrub( $text, 'UTF-8' );
	mb_substitute_character( $prev_replacement_character );

	return $scrubbed;
}

История изменений

ВерсияОписание
6.9.0 Introduced.

Что будем искать? Например,Продвижение

Этот сайт использует куки-файлы. Оставаясь на сайте, Вы соглашаетесь на их использование. Для получения дополнительной информации, пожалуйста, ознакомьтесь с политикой в отношении персональных данных.