跳转到内容
新建笔记

Python 基础操作:文本、切片、循环与数列

日常 Python 操作经常围绕文本、索引和循环展开。先确认数据类型,再区分“生成一个值”“修改对象”和“重复执行”,可以避开许多看似相近的写法。本文以 Python 3.11 为基线;最后的图像裁剪例子额外使用 NumPy。

数字文本、补零、清理与字典更新

跳转到“数字文本、补零、清理与字典更新”

int("000020") 解析十进制整数,str(20).zfill(6) 生成带前导零的文本。两者转换方向相反。strip("0") 则从两端移除字符 0,既不是数字转换,也不是只删除前导零:对原例 "000020",结果为 "2"。

import math
num_str = "000020"
num = int(num_str)
assert num == 20
assert str(num).zfill(6) == "000020"
assert num_str.strip("0") == "2"
assert num_str.lstrip("0") == "20"
assert num_str.rstrip("0") == "00002"
assert "000".lstrip("0") == "" and int("000") == 0
assert " hello world! \n".strip() == "hello world!"
assert "a-b-a".replace("a", "x") == "x-b-x"
# 原记录中的两项数值没有注明单位,这里只演示十进制文本转换。
raw_values = ["113.8029", "116.3988"]
values = [float(value) for value in raw_values]
assert math.isclose(values[1] - values[0], 2.5959)
dict1 = {"a": 1, "b": 2}
dict2 = {"c": 3, "d": 4}
result = dict1.update(dict2)
assert result is None
assert dict1 == {"a": 1, "b": 2, "c": 3, "d": 4}
dict1.update({"a": 9})
assert dict1["a"] == 9
print(num, str(num).zfill(6), dict1)

strip() 默认处理空格、制表符、换行等空白字符,不是只删除 ASCII 空格。lstrip()、rstrip() 分别处理左、右两端;replace() 替换匹配子串。处理数字时优先按数字格式解析,不要先随意删字符掩盖错误。方法边界见字符串,字典更新与复制见字典。

输出在一行、同时取得索引和值

跳转到“输出在一行、同时取得索引和值”

print 默认以换行结束,end=" " 改为一个空格;sep 则控制同一次调用中多个参数之间的分隔。下面把输出暂存到 StringIO,使空格和换行也能被准确检查。

from io import StringIO
output = StringIO()
for i in range(10):
print(i, end=" ", file=output)
assert output.getvalue() == "0 1 2 3 4 5 6 7 8 9 "
print(output.getvalue().rstrip())
colors = ["red", "green", "blue"]
pairs = list(enumerate(colors))
assert pairs == [(0, "red"), (1, "green"), (2, "blue")]
for index, color in pairs:
print(index, color)
assert list(enumerate(colors, start=1))[0] == (1, "red")
linemod_cls_names = ["ape", "benchvise", "camera"]
class_names = {cls_id: cls_name for cls_id, cls_name in enumerate(linemod_cls_names)}
assert class_names == {0: "ape", 1: "benchvise", 2: "camera"}

enumerate 产生计数与元素;名称写成 cls_id、cls_name 不会自动决定真实数据集的类别编号。这里仅展示原笔记的变量用途,示例列表是本程序定义的三项列表。若正式数据另有编号规则,应使用该规则,而不是直接假定从零连续编号。

切片:终点不包含,方向由步长决定

跳转到“切片:终点不包含,方向由步长决定”

通用写法为 sequence[start:stop:step]。第二个冒号用于引入步长;不是只有写成相邻的 :: 才能指定步长,a[1:8:2] 同样合法。步长不能为零。

正步长时省略的边界覆盖从左到右的范围,负步长时省略的边界适配从右到左。负索引是相对结尾的位置;负步长才决定反向遍历。不要把二者混为一谈,也不要把负步长的终点简单读成“到 stop-1”。

a = list(range(12))
assert a[1:5] == [1, 2, 3, 4]
assert a[:5] == [0, 1, 2, 3, 4]
assert a[3:] == list(range(3, 12))
assert a[1:5:2] == [1, 3]
assert a[1:8:2] == [1, 3, 5, 7]
assert a[::2] == [0, 2, 4, 6, 8, 10]
assert a[::3] == [0, 3, 6, 9]
assert a[::-1] == list(reversed(a))
assert a[-5:-2] == [7, 8, 9]
assert a[-6:10:2] == [6, 8]
tmp = [1, 2, 3, 4, 5, 6]
assert tmp[5::-2] == [6, 4, 2]
assert tmp[5:-1:-2] == []
assert a[100:200] == []
try:
a[::0]
except ValueError:
pass
else:
raise AssertionError("slice step cannot be zero")
original = [[1], [2]]
shallow = original[:]
assert shallow == original and shallow is not original
assert shallow[0] is original[0]
assert "abcdef"[::2] == "ace"
assert (1, 2, 3)[::-1] == (3, 2, 1)
print("slice checks passed")

tmp[5::-2] 省略了终点;tmp[5:-1:-2] 则明确把终点设为最后一个元素的位置,因此该例为空。切片越界通常被裁剪,而单个越界索引会抛出 IndexError。

列表的完整切片生成浅复制;嵌套成员仍可共享。不能从 sequence[:] 推出所有类型都会创建一个独立对象:不可变序列可能复用对象,NumPy 数组的基础切片通常是视图。详见列表复制关系和 Python 序列切片。

for、while 与 range 的实际执行次数

跳转到“for、while 与 range 的实际执行次数”

for 从可迭代对象逐项获取值,不是统一实现成“先赋值再比较上界”。可迭代对象可能没有 len(),也可能无限产生值;break、异常或迭代期间的数据变化都会影响循环次数。while 则在每轮开始前判断条件,通常需要明确地更新使条件最终为假的状态。

下面同时保留 range(1, 4)、步长为一的 range(2, 129) 和步长为二的 range(2, 129, 2):

seen = []
for item in [1, 2, 3, 4, 5]:
seen.append(item)
assert seen == [1, 2, 3, 4, 5]
count = 0
while count < 5:
count += 1
assert count == 5
for args, expected_count, expected_last in [
((1, 4), 3, 3),
((2, 129), 127, 128),
((2, 129, 2), 64, 128),
((5, -2, -2), 4, -1),
]:
sequence = range(*args)
visits = 0
for i in sequence:
visits += 1
assert visits == len(sequence) == expected_count
assert i == sequence[-1] == expected_last
visits = 0
for item in (n * n for n in range(100)):
visits += 1
if visits == 3:
break
assert visits == 3
i = "unchanged"
for i in range(3, 3):
raise AssertionError("empty range cannot enter loop")
assert i == "unchanged"
print("loop checks passed")

空循环不会为循环变量产生一个“最后的值”:若变量此前存在,它保持原值;若此前未绑定,之后直接读取会报错。非空、完整执行的 range(start, stop, step) 才可以按以下计数和最后一项公式分析。

用整数公式计算范围与数列

跳转到“用整数公式计算范围与数列”

令起点为 aa、终点为 bb、步长为 ss。正步长要求每项小于终点,负步长要求每项大于终点;方向不匹配时范围为空。若范围非空,项数和最后一项为:

N=⌈b−as⌉,xlast=a+(N−1)s.N=\left\lceil\frac{b-a}{s}\right\rceil,\qquad x_{\mathrm{last}}=a+(N-1)s.

这个公式的非空前提不能省略。实际程序可直接使用 len(range(...));对可能超过平台长度上限的大范围,可用纯整数计算项数,避免先做浮点除法再取整而丢失精度。

等差数列的第 nn 项及前 nn 项和为:

an=a1+(n−1)d,Sn=n(a1+an)2=n[2a1+(n−1)d]2.a_n=a_1+(n-1)d,\qquad S_n=\frac{n(a_1+a_n)}{2} =\frac{n[2a_1+(n-1)d]}{2}.

等比数列使用公比 rr:

an=a1rn−1,Sn={a11−rn1−r,r≠1,na1,r=1.a_n=a_1r^{n-1},\qquad S_n=\begin{cases} a_1\dfrac{1-r^n}{1-r}, & r\ne1,\\ na_1, & r=1. \end{cases}

这些项数公式针对正整数 nn。两类数列都是离散序列;不能把“等差数列的连续项”解释成连续变量。循环是执行机制,可以生成数列,也可以执行与数列无关的操作。

def range_count(start, stop, step=1):
if step == 0:
raise ValueError("step must not be zero")
if step > 0:
return max(0, (stop - start + step - 1) // step)
size = -step
return max(0, (start - stop + size - 1) // size)
for start in range(-4, 5):
for stop in range(-4, 5):
for step in [-3, -2, -1, 1, 2, 3]:
actual = range(start, stop, step)
n = range_count(start, stop, step)
assert n == len(actual)
if n:
assert start + (n - 1) * step == actual[-1]
assert range_count(2, 129, 2) == 64
assert range_count(0, 10**40, 3) == (10**40 + 2) // 3
arithmetic = list(range(2, 129, 2))
n, first, last = len(arithmetic), arithmetic[0], arithmetic[-1]
assert n == 64 and last == first + (n - 1) * 2
assert sum(arithmetic) == n * (first + last) // 2 == 4160
first, ratio, n = 3, 2, 4
geometric = [first * ratio**k for k in range(n)]
assert geometric == [3, 6, 12, 24]
assert geometric[-1] == first * ratio ** (n - 1)
assert sum(geometric) == first * (ratio**n - 1) // (ratio - 1) == 45
assert sum([3] * 4) == 4 * 3
print("range and sequence formula checks passed")

range_count 示例约定传入整数;这里的整数除法适用于示例中的整数组合,不能把等比求和公式里的除号在任意小数公比下都换成 //。len(range(...)) 也可能因长度超出平台可表示范围而报 OverflowError,这与范围本身无法表示是两件事。

向上、向下、截断与就近取整

跳转到“向上、向下、截断与就近取整”

math.ceil 向正无穷方向、math.floor 向负无穷方向,math.trunc 和有限浮点数的 int 转换向零截断。Python 内置 round 在精确中点采用“取偶数”,不是一律把 .5 远离零。

import math
from decimal import Decimal, ROUND_HALF_UP
assert math.ceil(3.1) == 4 and math.ceil(-3.7) == -3
assert math.floor(3.7) == 3 and math.floor(-3.1) == -4
assert math.trunc(3.9) == 3 and math.trunc(-3.9) == -3
assert int(3.9) == 3 and int(-3.9) == -3
assert round(3.5) == 4 and round(2.5) == 2
assert round(-3.5) == -4 and round(-2.5) == -2
assert round(2.4) == 2
assert round(2.675, 2) == 2.67
assert Decimal("2.5").quantize(Decimal("1"), rounding=ROUND_HALF_UP) == Decimal("3")
assert Decimal("-2.5").quantize(Decimal("1"), rounding=ROUND_HALF_UP) == Decimal("-3")
print("rounding checks passed")

2.675 的例子还涉及二进制浮点近似,并非只看十进制书写形式就能决定结果。若需求明确规定十进制舍入规则,用字符串构造 Decimal 并选择相应模式。依据:round、math 的取整函数。

类型注解与一个完整的裁剪函数

跳转到“类型注解与一个完整的裁剪函数”

函数注解语法和类型提示的发展有不同时间节点:函数注解在 Python 3.0 已有,标准类型提示方案见 Python 3.5 的 PEP 484,value: int = 5 这种变量注解语法由 Python 3.6 引入。注解本身不会强制运行时检查,原稿只写 pass 的 image_crop 也不会返回裁剪图像。

下面将边界明确为 (x0, y0, x1, y1),右端和下端不包含;要求非空且完全位于图像内,返回独立副本。代码使用 NumPy,可处理二维灰度图或带额外通道维度的数组。

from collections.abc import Sequence
import operator
import numpy as np
def image_crop(img: np.ndarray, bbox: Sequence[int]) -> np.ndarray:
if not isinstance(img, np.ndarray) or img.ndim < 2:
raise ValueError("img must be an array with at least two dimensions")
if len(bbox) != 4:
raise ValueError("bbox must contain x0, y0, x1, y1")
if any(isinstance(value, (bool, np.bool_)) for value in bbox):
raise TypeError("boolean coordinates are not accepted")
x0, y0, x1, y1 = (operator.index(value) for value in bbox)
height, width = img.shape[:2]
if not (0 <= x0 < x1 <= width and 0 <= y0 < y1 <= height):
raise ValueError("bbox must be nonempty and inside the image")
return img[y0:y1, x0:x1].copy()
value_name: int = 5
assert value_name == 5
image = np.arange(4 * 5 * 3).reshape(4, 5, 3)
cropped = image_crop(image, [1, 1, 4, 3])
assert cropped.shape == (2, 3, 3)
assert np.array_equal(cropped, image[1:3, 1:4])
assert not np.shares_memory(cropped, image)
cropped[0, 0, 0] = -1
assert image[1, 1, 0] != -1
try:
image_crop(image, [0, 0, 6, 3])
except ValueError:
pass
else:
raise AssertionError("out-of-bounds crop must be rejected")
print("annotated image crop checks passed")

这是为原接口补充的一种明确实现;需要裁剪越界框、归一化坐标或保留共享视图的程序应另定规则。原稿导入的 Union 没有在签名中使用,因而此例不导入它。注解详见类型注解,数组视图与复制见 NumPy 基础索引。

读取文本:保留哪部分空白要先决定

跳转到“读取文本:保留哪部分空白要先决定”

原笔记的“txt 读取”只有空代码块。下面补充一个完整示例:在临时目录创建 UTF-8 文件,然后逐行读取;离开临时目录上下文后清理这个演示文件。strip() 会删除两端所有空白,若缩进或尾部空格有意义,就只去掉行结束符。

from pathlib import Path
from tempfile import TemporaryDirectory
with TemporaryDirectory() as directory:
filename = Path(directory) / "sample.txt"
filename.write_text(" alpha \nbeta\n\n", encoding="utf-8")
with filename.open("r", encoding="utf-8") as stream:
rows = [line.rstrip("\r\n") for line in stream]
assert rows == [" alpha ", "beta", ""]
assert [row.strip() for row in rows] == ["alpha", "beta", ""]
assert " a b ".split(" ") == ["", "a", "", "b", ""]
assert " a b\t".split() == ["a", "b"]
print("text file checks passed")

split(" ") 只把单个空格当分隔符并保留相邻分隔造成的空字段;不传参数的 split() 才按连续空白分组。更多输入输出参数见 print、enumerate、range与 open。