mercury

mirror of https://github.com/Mercury-Language/mercury.git synced 2026-04-16 09:53:36 +00:00

Author	SHA1	Message	Date
Zoltan Somogyi	d23c4f74a3	Update the style of more tests.	2020-10-06 19:20:18 +11:00
Peter Wang	e28ff5bbe7	Define string.between_codepoints more precisely and fix bug. library/string.m: Define string.between_codepoints in terms of codepoint_offset. Fix behaviour in the case where Start < 0, End < 0, End > Start tests/hard_coded/string_codepoint.exp: tests/hard_coded/string_codepoint.exp2: tests/hard_coded/string_codepoint.m: Extend test case.	2019-10-24 12:31:29 +11:00
Peter Wang	adc00f7fc1	Merge branch 'version-14_01-branch'	2015-10-04 18:51:05 +11:00
Peter Wang	17f749241c	Fix string.between_codepoints range clamping. library/string.m: Fix string.between_codepoints so that the clamping of start/end points works as documented. tests/hard_coded/string_codepoint.m tests/hard_coded/string_codepoint.exp tests/hard_coded/string_codepoint.exp2 Improve test case. NEWS: Mention fix.	2015-10-04 14:46:32 +11:00
Zoltan Somogyi	33eb3028f5	Clean up the tests in half the test directories. tests/accumulator/.m: tests/analysis_/.m: tests/benchmarks/.m: tests/debugger/.{m,exp,inp}: tests/declarative_debugger/.{m,exp,inp}: tests/dppd/.m: tests/exceptions/.m: tests/general/.m: tests/grade_subdirs/.m: tests/hard_coded/*.m: Make these tests use four-space indentation, and ensure that each module is imported on its own line. (I intend to use the latter to figure out which subdirectories' tests can be executed in parallel.) These changes usually move code to different lines. For the debugger tests, specify the new line numbers in .inp files and expect them in .exp files.	2015-02-14 20:14:03 +11:00
Peter Wang	b1af59cb29	Deprecate string.substring, string.foldl_substring, etc. in favour of Branches: main Deprecate string.substring, string.foldl_substring, etc. in favour of new procedures named string.between, string.foldl_between, etc. The "between" procedures take a pair of [Start, End) endpoints instead of Start, Count arguments. The reasons for this change are: - the "between" procedures are more convenient - "between" should be unambiguous. You can guess that it takes an End argument instead of a Count argument without looking up the manual. - Count arguments necessarily counted code units, but when working with non-ASCII strings, almost the only way that the values would be arrived at is by substracting one end point from another. - it paves the way for a potential change to replace string offsets with an abstract type. We cannot do that if users regularly have to perform a subtraction between two offsets. library/string.m: Add string.foldl_between, string.foldl2_between, string.foldr_between, string.between. Deprecate the old substring names. Replace string.substring_by_codepoint by string.between_codepoints. compiler/elds_to_erlang.m: compiler/error_util.m: compiler/rbmm.live_variable_analysis.m: compiler/timestamp.m: extras/posix/samples/mdprof_cgid.m: library/bitmap.m: library/integer.m: library/lexer.m: library/parsing_utils.m: mdbcomp/trace_counts.m: profiler/demangle.m: Conform to changes. tests/general/Mercury.options: tests/general/string_foldl_substring.exp: tests/general/string_foldl_substring.m: tests/general/string_foldr_substring.exp: tests/general/string_foldr_substring.m: tests/hard_coded/Mercury.options: tests/hard_coded/string_substring.m: Test both between and substring procedures. tests/hard_coded/string_codepoint.exp: tests/hard_coded/string_codepoint.exp2: tests/hard_coded/string_codepoint.m: Update names in test outputs. NEWS: Announce the change.	2011-06-15 01:05:34 +00:00
Peter Wang	3788a9d6fb	Improve Unicode support. Branches: main Improve Unicode support. Declare that we use the Unicode character set, and UTF-8 or UTF-16 for the internal string representation (depending on the backend). User code may be written to those assumptions. Other external encodings can be supported in the future by translating to/from Unicode internally. The `char' type now represents a Unicode code point. NOTE: questions about how to handle unpaired surrogate code points, etc. have been left for later. library/char.m: Define a `char' to be a Unicode code point and extend ranges appropriately. Add predicates: to_utf8, to_utf16, is_surrogate, is_noncharacter. Update some documentation. library/io.m: Declare I/O predicates on text streams to read/write code points, not ambiguous "characters". Text files are expected to use UTF-8 encoding. Supporting other encodings is for future work. Update the C and Erlang implementations to understand UTF-8 encoding. Update Java and C# implementations to read/write code points (Mercury char) instead of UTF-16 code units. Add `may_not_duplicate' attributes to some foreign_procs. Improve Erlang implementations of seeking and getting the stream size. library/string.m: Declare the string representations, as described earlier. Distinguish between code units and code points everywhere. Existing functions and predicates which take offset and length arguments continue to take them in terms of code units. Add procedures: count_code_units, count_codepoints, codepoint_offset, to_code_unit_list, from_code_unit_list, index_next, unsafe_index_next, unsafe_prev_index, unsafe_index_code_unit, split_by_codepoint, left_by_codepoint, right_by_codepoint, substring_by_codepoint. Make index, index_det call error/1 if an illegal sequence is detected, as they already do for invalid offsets. Clarify that is_all_alpha, is_all_alnum_or_underscore, is_alnum_or_underscore only succeed for the ASCII characters under each of those categories. Clarify that whitespace stripping functions only strip whitespace characters in the ASCII range. Add comments about the future treatment of surrogate code points (not yet implemented). Use Mercury format implementation when necessary instead of `sprintf'. The %c specifier does not work for code points which require multi-byte representation. The field width modifier for %s only works if the string contains only single-byte code points. library/lexer.m: Conform to string encoding changes. Simplify code dealing with \uNNNN escapes now that encoding/decoding is handled by the string module. library/term_io.m: Allow code points above 126 directly in Mercury source. NOTE: \x and \o codes are treated as code points by this change. runtime/mercury_types.h: Redefine `MR_Char' to be `int' to hold a Unicode code point. `MR_String' has to be defined as a pointer to `char' instead of a pointer to `MR_Char'. Some C foreign code will be affected by this change. runtime/mercury_string.c: runtime/mercury_string.h: Add UTF-8 helper routines and macros. Make hash routines conform to type changes. compiler/c_util.m: Fix output_quoted_string_lang so that it correctly outputs non-ASCII characters for each of the target languages. Fix quote_char for non-ASCII characters. compiler/elds_to_erlang.m: Write out code points above 126 normally instead of using escape syntax. Conform to string encoding changes. compiler/mlds_to_cs.m: Change Mercury `char' to be represented by C# `int'. compiler/mlds_to_java.m: Change Mercury `char' to be represented by Java `int'. doc/reference_manual.texi: Uncomment description of \u and \U escapes in string literals. Update description of C# and Java representations for Mercury `char' which are now `int'. tests/debugger/tailrec1.m: Conform to renaming. tests/general/string_replace.exp: tests/general/string_replace.m: Test non-ASCII characters to string.replace. tests/general/string_test.exp: tests/general/string_test.m: Test non-ASCII characters to string.duplicate_char, string.pad_right, string.pad_left, string.format_table. tests/hard_coded/char_unicode.exp: tests/hard_coded/char_unicode.m: Add test for new procedures in `char' module. tests/hard_coded/contains_char_2.m: Test non-ASCII characters to string.contains_char. tests/hard_coded/nonascii.exp: tests/hard_coded/nonascii.m: tests/hard_coded/nonascii_gen.c: Add code points above 255 to this test case. Change test data encoding to UTF-8. tests/hard_coded/string_class.exp: tests/hard_coded/string_class.m: Add test case for string.is_alpha, etc. tests/hard_coded/string_codepoint.exp: tests/hard_coded/string_codepoint.exp2: tests/hard_coded/string_codepoint.m: Add test case for new string procedures dealing with code points. tests/hard_coded/string_first_char.exp: tests/hard_coded/string_first_char.m: Add test case for all modes of string.first_char. tests/hard_coded/string_hash.m: Don't use buggy random.random/5 predicate which can overflow on a large range (such as the range of code points). tests/hard_coded/string_presuffix.exp: tests/hard_coded/string_presuffix.m: Add test case for string.prefix, string.suffix, etc. tests/hard_coded/string_set_char.m: Test non-ASCII characters to string.set_char. tests/hard_coded/string_strip.exp: tests/hard_coded/string_strip.m: Test non-ASCII characters to string stripping procedures. tests/hard_coded/string_sub_string_search.m: Test non-ASCII characters to string.sub_string_search. tests/hard_coded/unicode_test.exp: Update expected output due to change of behaviour of `string.to_char_list'. tests/hard_coded/unicode_test.m: Test non-ASCII character in separator string argument to string.join_list. tests/hard_coded/utf8_io.exp: tests/hard_coded/utf8_io.m: Add tests for UTF-8 I/O. tests/hard_coded/words_separator.exp: tests/hard_coded/words_separator.m: Add test case for `string.words_separator'. tests/hard_coded/Mmakefile: Add new test cases. Make special_char test case run on all backends. tests/hard_coded/special_char.exp: tests/valid/mercury_java_parser_follow_code_bug.m: Reencode these files in UTF-8. NEWS: Add a news entry.	2011-04-04 07:10:42 +00:00

7 Commits