Prerequisites
Background / Description
Read builds its numbered lines with str.splitlines(keepends=True) (src/agentscope/tool/_builtin/_read.py:522). After _normalize_newlines, splitlines() still breaks lines at \v, \f, \x1c, \x1d, \x1e, \x85, \u2028 and \u2029. Grep (rg -n), sed -n, wc -l and editors count only \n.
So in a file that contains one of those characters, every line after it has a different number in Read than in Grep:
- A form feed is a common page break in older Python, C and Emacs-style sources.
Grep reports 5: return 2 # needle, while Read(offset=5) returns def second():.
- A
\u2028 inside a string literal is legal in JS since ES2019 and also shows up in pasted text. Read shows that one line as two numbered lines. If the agent then copies those lines into Edit's old_string, joining them with \n, Edit answers old_string not found.
Expected: Read numbers lines the way Grep and the shell do, so an offset taken from Grep output, or a line copied from Read output, refers to the same text.
Write's line count (#2734) and the read cache that Write and Edit refresh use splitlines() too (_write.py:304,309, _edit.py:412), so they would need the same rule to stay consistent with Read.
Error Messages
Output of the script below on main (the form feed and \u2028 are shown escaped):
wc -l: 5
Grep: 5: return 2 # needle
sed -n 5p: return 2 # needle
Read offset=5: ' 5\tdef second():'
Read whole file:
1 def first():
2 return 1
3 \f
4
5 def second():
6 return 2 # needle
Read strings.js:
1 const msg = "first\u2028
2 second";
3 const other = 1;
Edit with the lines as Read showed them: error | Error: old_string not found in /tmp/tmpXXXX/strings.js
Steps to Reproduce
- Code:
import asyncio
import os
import subprocess
import tempfile
from agentscope.tool import Edit, Grep, Read
async def main() -> None:
with tempfile.TemporaryDirectory() as tmp:
path = os.path.join(tmp, "report.py")
# A form feed (page break) between two sections.
content = "def first():\n return 1\n\f\ndef second():\n return 2 # needle\n"
with open(path, "w", encoding="utf-8", newline="") as f:
f.write(content)
print("wc -l:", subprocess.run(["wc", "-l", path], capture_output=True, text=True).stdout.split()[0])
grep = await Grep()(pattern="needle", path=path, output_mode="content", n=True)
print("Grep:", grep.content[0].text.strip().splitlines()[-1])
print("sed -n 5p:", subprocess.run(["sed", "-n", "5p", path], capture_output=True, text=True).stdout.rstrip())
read5 = await Read()(file_path=path, offset=5, limit=1)
print("Read offset=5:", repr(read5.content[0].text))
read_all = await Read()(file_path=path)
print("Read whole file:")
print(read_all.content[0].text)
# A line holding U+2028 inside a string literal.
path2 = os.path.join(tmp, "strings.js")
with open(path2, "w", encoding="utf-8", newline="") as f:
f.write('const msg = "first\u2028second";\nconst other = 1;\n')
read_js = await Read()(file_path=path2)
print("Read strings.js:")
print(read_js.content[0].text)
edit = await Edit()(file_path=path2, old_string='const msg = "first\nsecond";', new_string='const msg = "x";')
print("Edit with the lines as Read showed them:", edit.state, "|", edit.content[0].text[:80])
asyncio.run(main())
- Run:
python repro.py
- See the output above.
A possible fix would be a small helper next to _normalize_newlines that splits only on \n and keeps the line endings. Read, Write and Edit would use it in place of splitlines(). The existing CRLF/CR normalization would stay as it is. I'm happy to open a PR with tests if this direction works for you.
Environment
- AgentScope Version: 2.0.10dev (
main at a833ab9)
- Python Version: 3.11.15
- OS: Linux
Prerequisites
Background / Description
Readbuilds its numbered lines withstr.splitlines(keepends=True)(src/agentscope/tool/_builtin/_read.py:522). After_normalize_newlines,splitlines()still breaks lines at\v,\f,\x1c,\x1d,\x1e,\x85,\u2028and\u2029.Grep(rg -n),sed -n,wc -land editors count only\n.So in a file that contains one of those characters, every line after it has a different number in
Readthan inGrep:Grepreports5: return 2 # needle, whileRead(offset=5)returnsdef second():.\u2028inside a string literal is legal in JS since ES2019 and also shows up in pasted text.Readshows that one line as two numbered lines. If the agent then copies those lines intoEdit'sold_string, joining them with\n,Editanswersold_string not found.Expected:
Readnumbers lines the wayGrepand the shell do, so an offset taken fromGrepoutput, or a line copied fromReadoutput, refers to the same text.Write's line count (#2734) and the read cache thatWriteandEditrefresh usesplitlines()too (_write.py:304,309,_edit.py:412), so they would need the same rule to stay consistent withRead.Error Messages
Output of the script below on
main(the form feed and\u2028are shown escaped):Steps to Reproduce
python repro.pyA possible fix would be a small helper next to
_normalize_newlinesthat splits only on\nand keeps the line endings.Read,WriteandEditwould use it in place ofsplitlines(). The existing CRLF/CR normalization would stay as it is. I'm happy to open a PR with tests if this direction works for you.Environment
mainat a833ab9)